System
A system using user profile information and generative AI generates interactive documentaries, allowing for personalized and dynamic content based on user selections and feedback, addressing the challenge of limited quality documentaries and viewer boredom.
Patent Information
- Application Number
- JP2024137441
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Producing documentaries requires specialized knowledge and talent, limiting the number of quality documentaries and making it difficult for viewers to find content that matches their interests, leading to boredom with repetitive content.
A system that includes acquiring user profile information, generating documentary content using a generative AI model, providing interactive questions, regenerating content based on user selections, indexing and storing content, and utilizing user feedback for personalized and dynamic content generation.
Enables an infinite number of customized documentary experiences tailored to individual interests and emotional states, reducing viewer boredom and maximizing learning outcomes.
Smart Images

Figure 2026034320000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Producing a documentary requires specialized knowledge and a great deal of talent, and the number of quality documentaries is limited. This makes it difficult for viewers to find documentaries that match their interests and knowledge level. Furthermore, viewers often become bored with a single documentary and require continuous new content. There is a need for an interactive platform that can solve these challenges and allow users to experience an infinite number of new stories. [Means for solving the problem]
[0005] The present invention solves these problems by providing a system that includes: means for acquiring user profile information; means for generating documentary content based on the user's interests and selections using a generative AI model; means for providing the generated content to the user and generating and presenting interactive questions; means for regenerating subsequent documentary content based on the user's selections; and means for indexing and storing the generated content in a database. The present invention also provides a more personalized experience by further including means for collecting feedback from the user's device and utilizing the feedback in subsequent content generation. Furthermore, the present invention also includes means for saving the generated documentary content in a reusable form and reusing it to accommodate other user selections, thereby reducing computational costs and providing content more efficiently.
[0006] "User profile information" is attribute information about individual users, such as their interests, preferences, and past viewing history.
[0007] A "generative AI model" is a software model that uses machine learning algorithms to generate new content based on specified input data.
[0008] "Documentary content" refers to videos that are produced based on real-life information or footage and are intended to be educational or informational.
[0009] An "interactive question" is a question-style interface that allows the viewer to make choices that affect the content that is subsequently generated.
[0010] A "terminal" is an electronic device that a user uses to view interactive content and to make inputs.
[0011] "Server" means a central processing unit that receives requests from user terminals and generates or provides content using generative AI models.
[0012] A "next chapter" is a new documentary video segment that continues or completes the story, generated based on the user's choices and interests.
[0013] "Indexing" is the process of organizing generated content and storing it in a database so that it can be easily searched and reused.
[0014] "Feedback" refers to reactions such as ratings and comments provided by users after viewing, and is information used to customize future content.
[0015] "Reusable form" refers to a method of arrangement in which the generated content is stored so that it can be used by other users and quickly provided when needed.
[0016] An "interface" is a screen or method of operation that allows a user to interact with a system. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0039] Server Operation
[0040] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0041] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0042] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0043] Device behavior
[0044] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0045] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0046] User Actions
[0047] Users log in to the platform and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that appear, which determine the content of the next video. After watching, users provide feedback that will be reflected in the next video. Through this interactive selection and feedback process, users can experience infinitely customized documentary content.
[0048] Specific examples
[0049] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0050] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0051] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0052] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] A user logs into the platform. The user enters their authentication information on the login screen and is authenticated.
[0056] Step 2:
[0057] The terminal sends the user's login information to the server, which then retrieves the user's profile information from the database based on the received authentication information.
[0058] Step 3:
[0059] The server extracts the user's interests and past viewing history from their profile information and uses that information to send an initial documentary content generation request to the generative AI model.
[0060] Step 4:
[0061] The generative AI model generates an initial documentary video and returns it to the server, which then receives the video file and sends it to the device.
[0062] Step 5:
[0063] The terminal plays the initial documentary video and provides a viewing interface to the user, who then watches the video.
[0064] Step 6:
[0065] The server generates interactive questions based on the video content and sends the questions and options to the device, which then displays them to the user.
[0066] Step 7:
[0067] The user selects one of the options for the question presented, and the selection information is sent from the terminal to the server.
[0068] Step 8:
[0069] The server receives the user's selection and sends another request to the generative AI model to generate the next documentary chapter, updating the profile information if necessary.
[0070] Step 9:
[0071] The server receives the generated video file of the next documentary chapter and sends it to the device, which plays this new video.
[0072] Step 10:
[0073] The user watches a new documentary chapter and makes selections in response to the interactive questions that are presented again, and the process is repeated.
[0074] Step 11:
[0075] The server indexes and stores each generated video in a database so that it can be quickly provided to other users who make similar selections.
[0076] Step 12:
[0077] After watching a video, the user inputs feedback into the device, which then sends the feedback to the server and uses it to generate the next content.
[0078] This allows users to experience an infinite amount of customized documentary content.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] Conventional video content is difficult to customize according to user interests, and viewers often become bored with the same content. Furthermore, while systems existed that provided interactive experiences, they lacked the ability to dynamically generate content based on user choices. As a result, the appeal of the viewing experience was limited, and learning benefits were reduced.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes means for acquiring user profile information, means for generating interactive content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selections, means for indexing and storing each generated chapter in a database, and means for reusing the generated content to quickly provide it to other users who make the same selections, thereby enabling dynamic and infinite generation of customized documentary content according to the user's interests.
[0084] "User profile information" is data that indicates the individual attributes and behavioral patterns of each user, such as the user's interests, past viewing history, and selection history.
[0085] "Generative AI model" refers to an artificial intelligence algorithm or framework for dynamically generating documentary content based on user profile information and selection information.
[0086] An "interactive question" is a question that is displayed while the user is watching, prompting the user to make a choice, and is an element that influences the generation of the next content.
[0087] "Indexing" is the process of organizing and classifying generated content into a database so that it can be efficiently searched and referenced.
[0088] "Reuse" refers to reducing the computational cost of the system by making content once generated available to other users who make the same or similar choices.
[0089] "Feedback" refers to ratings and comments provided by users after viewing a content, and is sent to the server as a reference for the next content generation.
[0090] "Terminal" refers to an apparatus or device through which a user views content and inputs selections and feedback.
[0091] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0092] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history, which serves as the basic data for content generation. The server then generates documentary content using a generative AI model (for example, OpenAI's GPT-4 (registered trademark) or TENSORFLOW (registered trademark)). The prompt uses a format such as "Generate the content of the first interactive documentary based on the user's profile information." The generated content is then sent to the device.
[0093] The user uses a device to watch this initial documentary content. While watching, the server generates interactive questions, sends them to the device, and presents them to the user. For example, in response to the question, "Which part are you interested in next?", options such as "Pyramids," "Pharaoh's History," and "Daily Life" are presented. The user's selection is sent from the device to the server, and the next content is generated based on this. Again, the prompt is in the form, "Based on the user's selection and profile information, generate the next chapter of the documentary about the pyramids."
[0094] Each generated chapter is stored and indexed in a database on the server, allowing it to be quickly provided to other users who make the same selection, reducing computational costs. The content generated through reuse also corresponds to other users' selections.
[0095] After watching, users can enter their feedback into their device, which then sends it to the server. The server then incorporates this feedback into future content generation. For example, if a user gives a high rating to the "Pyramid" chapter, the server will take this into account in future content generation and adjust the content generation to provide more detailed information.
[0096] As a concrete example, if a user is interested in "history," the server generates a documentary on the theme of "ancient Egypt" based on the user's profile information. The prompt text used is, "Based on the user's profile information, generate an interactive documentary about ancient Egypt." The generated video is displayed on the device, and the user begins watching. Interactive questions are presented during viewing, and the next documentary chapter is generated based on the user's selection. In this way, an infinite documentary experience completely customized to the user's interests is realized.
[0097] This system allows users to enjoy an interactive and dynamic viewing experience by providing documentary content customized to their interests, thereby preventing user boredom and maximizing learning effectiveness.
[0098] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0099] Step 1:
[0100] The server receives the user's identification information (user ID) when the user logs into the platform. As input, the user ID is sent to the server. The server uses this user ID to retrieve the corresponding profile information (interests, viewing history, etc.) from the database. As output, the retrieved profile information is ready to be passed to the generative AI model.
[0101] Step 2:
[0102] The server uses the acquired profile information to create a prompt appropriate for the generative AI model. This prompt includes instructions for generating initial documentary content based on the user's interests. For example, a prompt might be created such as, "Generate an interactive documentary about ancient Egypt based on the user's profile information." The acquired profile information is used as input, and the prompt is generated as output.
[0103] Step 3:
[0104] The server inputs the generated prompt sentences into a generative AI model to generate initial documentary content. The generative AI model generates data such as text, images, and videos based on the input prompt sentences. The output is the initial documentary content, which includes content that is in line with the user's interests.
[0105] Step 4:
[0106] The server sends the generated initial documentary content to the terminal, which uses the generated documentary content as input and transfers it to the terminal as output, in an appropriate format (e.g., video file or streaming data).
[0107] Step 5:
[0108] The terminal displays the initial documentary content received from the server to the user, who then begins viewing the content, using the documentary content received from the server as input and providing viewing to the user as output.
[0109] Step 6:
[0110] The server generates interactive questions according to the viewing progress. For example, when a specific scene in a video is reached, the server generates a question such as "Which part are you interested in next?" and presents options such as "Pyramids," "Pharaoh's history," and "Daily life." The viewing progress data is used as input, and the interactive questions are generated as output.
[0111] Step 7:
[0112] The terminal presents the user with interactive questions sent from the server. The user selects one of the options presented. The user's option is entered into the terminal as input, and the selection is sent to the server as output.
[0113] Step 8:
[0114] The server sends a prompt to the generative AI model again based on the user's selection and profile information to generate the next chapter of documentary content. The prompt might be something like, "Generate the next chapter of a documentary about pyramids based on the user's selection and profile information." The user's selection and profile information are used as input, and the documentary content for the next chapter is generated as output.
[0115] Step 9:
[0116] The server sends the generated documentary content of the next chapter to the terminal. The terminal receives this new content and provides it to the user. The generated documentary content is used as input and sent to the terminal as output.
[0117] Step 10:
[0118] A user watches new documentary content and then provides feedback. As input, the feedback information is input to the terminal, and as output, the feedback is sent to the server.
[0119] Step 11:
[0120] The server stores the feedback received from the device in a database and uses it for the next content generation. The user's feedback information is used as input and stored in the database as output. This feedback information is used to adjust the prompt text to be generated next time.
[0121] Step 12:
[0122] The server analyzes the accumulated feedback and reflects it in the next content generation. This allows for more highly customized content based on the user's interests and satisfaction. The feedback data is used as input, and the output is used to adjust the generated prompts and the generative AI model.
[0123] (Application example 1)
[0124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0125] Conventional content delivery systems have difficulty generating and providing personalized content based on user interests and preferences, resulting in inconsistent viewing experiences. Furthermore, the generated content lacks interactivity, making it difficult to maintain user interest over the long term. Furthermore, there is a lack of a way to utilize user feedback to improve the quality of the viewing experience.
[0126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0127] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, and means for delivering the generated content via a smartphone application and collecting user feedback, thereby enabling the generation of personalized content based on the user's interests, providing an interactive experience, and utilizing the collected feedback for future content generation.
[0128] "User profile information" is information including the user's interests, past viewing history, personal data, etc., and is the basis for content generation.
[0129] A "generative AI model" is a model that uses artificial intelligence to generate content based on user profile information and interactive choices.
[0130] "Documentary content" refers to content that is comprised of fact-based video and audio, and includes personalized content tailored to the user's interests.
[0131] An "interactive question" is a choice or question presented to generate different content based on a user's selection.
[0132] A "smartphone application" is an application that runs on a smartphone and provides an interface for users to answer interactive questions and view content.
[0133] "Feedback" refers to the act of a user providing opinions and impressions about content they have viewed, and the information provided can be used to help generate the next piece of content.
[0134] "Indexing" is the process of organizing generated content and storing it in a database for easy searching and reuse.
[0135] A system for implementing this invention acquires user profile information and uses that information to generate interactive documentary content based on the user's interests and preferences using a generative AI model. The system is composed of the following components:
[0136] 1. Server Operation
[0137] The server retrieves the user's profile information from a database. This information includes the user's interests and past viewing history. Based on this profile information, the server provides appropriate input data to a generative AI model to generate initial documentary content. The generative AI model used uses OpenAI GPT-3 (registered trademark), a general artificial intelligence platform. The generated content is then sent to the device, where the user can begin viewing.
[0138] 2. Device Operation
[0139] The user's smartphone application receives the initial documentary content from the server. When the user begins watching, the application displays interactive questions. The user's selections in response to these questions are sent from the device to the server. To generate the next documentary chapter, the server again invokes the generative AI model, inputting the user's selections and profile information. The newly generated documentary chapter is again sent to the device and served to the user.
[0140] 3. User Actions
[0141] Users log in to the smartphone application and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that determine the content of the next video. After watching, users provide feedback, and that information is reflected in the next content generation.
[0142] Specific examples
[0143] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0144] While watching, the question "Which part are you interested in next?" is displayed, and the options are "Pyramids," "Pharaoh's History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focusing on "Pyramids." Next, the following sentence is used as an example of a prompt sentence for the generative AI model:
[0145] Example prompt sentence:
[0146] "Username is interested in history, particularly ancient Egypt. Please generate a 10-minute documentary on the subject."
[0147] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0148] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0149] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0150] Step 1:
[0151] The server retrieves user profile information from a database. The input is a user ID or other identifying information, and the output is profile information including the user's interests and past viewing history. It filters and retrieves the required data from the database (e.g., MySQL®) using SQL queries.
[0152] Step 2:
[0153] The server generates a content generation prompt for the generative AI model based on the acquired profile information. The input is the user profile information, and the output is a prompt sentence for the generative AI model. Specifically, the server analyzes the profile information using a Python script and constructs an appropriate prompt sentence.
[0154] Step 3:
[0155] The server inputs a prompt into a generative AI model (e.g., OpenAI GPT-3) to generate the initial documentary content. The input is the prompt, and the output is the generated documentary content. Specifically, the server makes an API call and receives the generated text and video data.
[0156] Step 4:
[0157] The server sends the generated content to the user's smartphone app. The input is the generated documentary content, and the output is data transmission to the smartphone app. Data is sent using an HTTP request.
[0158] Step 5:
[0159] The device (smartphone app) displays the received documentary content to the user. The user begins watching and is presented with interactive questions. The input is the content received from the server, and the output is the interface displayed to the user. Specifically, the display uses UI elements created with React Native or Flutter (registered trademark).
[0160] Step 6:
[0161] The user answers the interactive questions presented to them. The input is the interactive question, and the output is the user's answer. When the user presses the selection button, the selection is registered within the application.
[0162] Step 7:
[0163] The terminal sends the user's answer information to the server. The input is the user's answer information, and the output is data sent to the server. The data is sent using an HTTP POST request.
[0164] Step 8:
[0165] The server uses the user's selections to generate the next documentary chapter. The input is the user's selections and profile information, and the output is a new prompt for the generative AI model and new generated content. Another API call is made to retrieve the necessary data.
[0166] Step 9:
[0167] The generated new content is then sent to the device again, and the user begins viewing it. The input is the generated new content, and the output is data sent to the smartphone app.
[0168] Step 10:
[0169] The user views new content and answers the interactive questions again. This process is repeated between the server, the device, and the user. The input is the new interactive questions, and the output is the user's new answers.
[0170] This allows users to enjoy an endless variety of personalized documentary experiences.
[0171] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0172] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[0173] Server Operation
[0174] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0175] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0176] Furthermore, the emotion engine uses the user's emotional data to tailor the next content or question. This emotional data is processed in real time, so the content provided is most appropriate for the user's current emotional state.
[0177] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0178] Device behavior
[0179] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0180] The device is equipped with an emotion engine that analyzes the user's emotions in real time from their facial expressions, voice, etc. This emotion data is sent to the server and used to generate the next content.
[0181] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0182] User Actions
[0183] Users log in to the platform and watch the initial documentary video provided. They make their own choices in response to interactive questions that appear while watching. The emotion engine also analyzes the user's emotions, and the emotional information is fed back to the system without any special operation.
[0184] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[0185] Specific examples
[0186] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0187] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0188] Additionally, the emotional state of the user while watching is analyzed by the emotion engine, and if the user is excited, for example, more detailed and visually stimulating content is provided.
[0189] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to their interests and emotions.
[0190] The system customizes documentaries according to individual users' interests and emotional state, providing an interactive viewing experience that combats viewer boredom and maximizes learning outcomes.
[0191] The processing flow will be explained below.
[0192] Server Processing
[0193] Step 1:
[0194] The server receives the user's login request and verifies the credentials.
[0195] Step 2:
[0196] The server retrieves the user's profile information from a database, which includes the user's interests, past viewing history, and individual viewing habits.
[0197] Step 3:
[0198] The server provides profile information to the generative AI model to generate the initial documentary content, which then generates the initial video based on this information.
[0199] Step 4:
[0200] The server receives the generated initial documentary video and transmits it to the terminal.
[0201] Step 5:
[0202] The server also receives data from the emotion engine to capture the user's current emotional state during video playback. The emotion data is processed in real time to influence the next content generation.
[0203] Terminal handling
[0204] Step 1:
[0205] The terminal displays an interface for the user to enter login information, which is then sent to the server.
[0206] Step 2:
[0207] The terminal plays back the initial documentary video received from the server and provides it to the user.
[0208] Step 3:
[0209] The emotion engine built into the device analyzes the user's emotions in real time from their facial expressions and voice, and this emotion data is sent to the server.
[0210] Step 4:
[0211] Interactive questions are displayed during the video, which contain multiple options for the user to select.
[0212] Step 5:
[0213] The user's selection results and emotion data are sent to the server and used as data for generating new content.
[0214] User Action
[0215] Step 1:
[0216] Users log in to the platform and watch the initial video.
[0217] Step 2:
[0218] While watching, interactive questions are displayed and the user is prompted to select one of the options.
[0219] Step 3:
[0220] The emotion engine analyzes the user's emotions in real time and sends the data to the server.
[0221] Step 4:
[0222] The user then watches the next newly generated video and answers the questions again.
[0223] Step 5:
[0224] After viewing, the user enters feedback, and this data is also sent to the server.
[0225] Specific examples
[0226] Step 1:
[0227] When a user logs in with an interest in "history," the server obtains this interest information.
[0228] Step 2:
[0229] The server requests an initial documentary related to "Ancient Egypt" from the generative AI model, while the device detects the user's emotional state, analyzing facial expressions and voice sounds that indicate excitement or curiosity in real time.
[0230] Step 3:
[0231] An initial video is generated and provided to the user. While watching the video, the user is prompted with the question, "Which part are you interested in next?" The options offered are "Pyramids," "Pharaoh's History," and "Everyday Life."
[0232] Step 4:
[0233] When the user selects "Pyramid" and the emotion engine detects the user's state of excitement, the server adjusts the next video it generates to be more visually stimulating.
[0234] Step 5:
[0235] A new video is generated and sent to the user's device, the user starts watching again, and the same process continues.
[0236] This system allows users to experience a wide variety of documentary content optimized for their emotional state.
[0237] Example 2
[0238] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0239] While conventional documentary content generation systems have the ability to generate content based on user interests and selection, they do not fully utilize user emotions or real-time feedback. This makes it difficult to fine-tune personalization based on individual user interests and emotional states, resulting in a poor viewing experience. Furthermore, it is difficult to efficiently reuse generated content, resulting in a waste of computing resources.
[0240] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0241] In this invention, the server includes a means for acquiring user profile information, a means for generating documentary content based on the user's interests and selections using a generative AI model, a means for collecting emotion data in real time using an emotion engine that recognizes the user's emotions and using that data to generate subsequent content, and a means for indexing and storing the generated content in a database. This makes it possible to provide highly personalized documentary content based on the user's interests and emotions, thereby realizing efficient use of computing resources and an improved quality of viewing experience.
[0242] "User profile information" is information that includes personal data such as a user's interests, past viewing history, preferences, and the like.
[0243] A "generative AI model" is an artificial intelligence system that generates appropriate documentary content based on a user's profile information and choices.
[0244] "Documentary content" is content consisting of video, audio and text based on facts and information.
[0245] An "interactive question" is a question with options that is presented to the user while they are watching, and the user's choice is reflected in the next content.
[0246] The "emotion engine" is a technology that collects and analyzes emotional data from users' facial expressions, voice, etc. in real time.
[0247] "Emotion data" is data that indicates the user's current emotional state, and is used to generate the next content.
[0248] A "database" is a system that efficiently stores and manages acquired information and generated content.
[0249] "Indexing" is the process of organizing data into a searchable format that makes queries within a database more efficient.
[0250] "Feedback" refers to information such as impressions and evaluations provided by users after viewing a content, and is useful for creating the next content.
[0251] "Reuse" refers to the reuse of content that has been generated once in a manner that corresponds to the selections of other users.
[0252] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[0253] Server Operation
[0254] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history. The server then provides this information to a generative AI model to generate the initial documentary content. The generative AI model can be OpenAI GPT-4 or similar.
[0255] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The generative AI model is again invoked to generate the next documentary chapter based on the user's choices, inputting their selections and profile information.
[0256] Furthermore, the system uses the user's emotional data recognized by the emotion engine to tailor the next content or question. Based on the emotional data processed in real time, the system provides content that best suits the user's current emotional state. Each generated chapter is stored and indexed in a database, allowing it to be quickly provided to other users who make the same selection, reducing computational costs.
[0257] Device behavior
[0258] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0259] The device also has a built-in emotion engine that analyzes emotions from the user's facial expressions and voice in real time and sends the data to the server. This emotion data is used to generate new content. The newly generated documentary chapter is sent to the device and provided to the user. Feedback entered by the user after watching the video is also sent from the device to the server and used to generate the next piece of content.
[0260] User Actions
[0261] Users log in to the platform and watch the initial documentary video provided to them. They then select options in response to interactive questions that appear while they are watching. Furthermore, emotional information analyzed by the emotion engine is fed back to the system in real time, so no special operations are required.
[0262] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[0263] Specific examples
[0264] For example, if a user is interested in "history," the server uses this interest information to request the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary. While watching, the question "Which part are you interested in next?" is displayed, and the options are presented: "Pyramids," "Pharaoh's history," and "Daily life." If the user selects "Pyramids," the selection is sent to the server, and the next chapter of the documentary is generated with content focused on "Pyramids."
[0265] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[0266] Prompt Sentence Examples
[0267] "If a user is interested in history, we can have the generative AI model generate a documentary on the theme of 'Ancient Egypt,' taking into account the user's specific profile information and past viewing history to provide interesting content."
[0268] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0269] Step 1:
[0270] A user logs in to the platform. The user accesses the platform and enters their username and password on the login screen. When they press the login button, the information is sent to the server. The server performs authentication, and if successful, starts the user's session.
[0271] Input: Username, Password
[0272] Output: User authentication result (success / failure), session start
[0273] Step 2:
[0274] The server retrieves the user's profile information: After successful authentication, the server retrieves the user's profile information (interests, past viewing history, etc.) from the database. This profile information is used in the next processing step.
[0275] Input: User ID
[0276] Output: User profile information
[0277] Step 3:
[0278] The server provides prompts to the generative AI model to generate initial documentary content. The server creates appropriate prompt sentences for the generative AI model based on the profile information and sends them to the model. The generative AI model generates documentary content based on these prompt sentences.
[0279] Input: User profile information, prompt text
[0280] Output: Initial documentary content
[0281] Step 4:
[0282] The server sends the generated documentary content to the device. The server receives the content generated by the generative AI model and sends it to the user's device. The content is ready for the user to watch.
[0283] Input: Early documentary content
[0284] Output: Content delivery to devices
[0285] Step 5:
[0286] The terminal displays the documentary content and the user starts viewing it. The terminal plays the content received from the server and makes it available for viewing by the user. The user starts viewing the content.
[0287] Input: Documentary content from the server
[0288] Output: Playback begins, user watches
[0289] Step 6:
[0290] The terminal displays interactive questions. While the user is watching the documentary, the terminal displays the interactive questions received from the server and presents the user with options.
[0291] Input: Interactive questions from the server
[0292] Output: Screen presenting options to the user
[0293] Step 7:
[0294] The user selects an option for an interactive question. The user selects one option from multiple options and enters it into the terminal. The terminal sends the selected information to the server.
[0295] Input: User's choice
[0296] Output: Sending choices from the terminal to the server
[0297] Step 8:
[0298] The terminal transmits the selection to the server. After receiving the user's selection, the terminal transmits the data to the server, providing input data for the next content generation.
[0299] Input: User's choice
[0300] Output: Send selected data to the server
[0301] Step 9:
[0302] The server requests a new documentary chapter from the generative AI model. The server generates a prompt for the new documentary chapter based on the user's choices and sends it to the generative AI model. The generative AI model then generates the new chapter based on that prompt.
[0303] Input: User choices, prompt text
[0304] Output: New documentary chapter
[0305] Step 10:
[0306] The device collects the user's emotional data and sends it to the server. The device's built-in emotion engine analyzes the user's emotional data while watching and sends the information to the server.
[0307] Input: User's emotional state
[0308] Output: Send emotion data to the server
[0309] Step 11:
[0310] The server generates the next content based on the emotion data. It analyzes the received emotion data and adjusts the content and questions of the next documentary chapter. The generated documentary chapter is indexed and stored in a database.
[0311] Input: User emotion data
[0312] Output: adjusted documentary chapters, stored in a database
[0313] Step 12:
[0314] The server sends new documentary chapters to the terminal, and the server sends the new chapters to the terminal so that the user can continue watching.
[0315] Enter: a new documentary chapter
[0316] Output: New documentary streaming to your device
[0317] Step 13:
[0318] User continues watching. The user watches the newly received documentary chapter and continues their viewing experience.
[0319] Enter: a new documentary chapter
[0320] Output: User continues watching
[0321] Step 14:
[0322] Users provide feedback on videos. After watching, users input their impressions and ratings into their devices, and the feedback is sent to the server. This information is used to generate new content for the next time.
[0323] Input: User feedback
[0324] Output: Send feedback to the server
[0325] The above processing steps enable the system of the present invention to provide highly personalized documentary content according to the user's interests, preferences, and emotional state.
[0326] (Application example 2)
[0327] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0328] Conventional documentary content generation systems personalize content based on user interests and selections, but they fail to take into account the user's emotional state. This makes it difficult to provide content that adapts to the user's emotional state, and it is therefore difficult to significantly improve the satisfaction of the viewing experience. Furthermore, user feedback is often not adequately reflected in the next content generation, limiting the improvement of content quality.
[0329] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0330] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selection, means for indexing and storing the generated content in a database, and means for recognizing the user's emotions using an emotion engine and reflecting the emotion data in the generation of the next content, thereby enabling the provision of sophisticated personalized content according to the user's emotional state.
[0331] "User profile information" is individual information about a user, such as the user's interests and past viewing history.
[0332] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate documentary content based on user interests and choices.
[0333] An "interactive question" is a question that is presented to a user to elicit a choice from the user.
[0334] The "emotion engine" is a system that analyzes the user's facial expressions and voice to obtain emotional data.
[0335] "Documentary content" refers to fact-based video works and informational content.
[0336] "Feedback" refers to data such as reactions, evaluations, and opinions from users.
[0337] "Personalization" is a method of providing individualized experiences and content.
[0338] A "database" is a system in which information is systematically stored and can be retrieved as needed.
[0339] "Indexing" is the process of organizing information in a database so that it can be quickly searched.
[0340] The system of the present invention acquires user profile information and generates interactive documentary content using a generative AI model based on that information. By combining it with an emotion engine, it is possible to provide more sophisticated personalized content.
[0341] Server Operation
[0342] The server first retrieves the user's profile information from a database, including the user's interests and past viewing history. The server then uses this profile information to provide appropriate input data to a generative AI model to generate initial documentary content. The generated documentary content is then sent to the device, where the user can begin viewing.
[0343] Interactive questions are generated and presented to the user as they watch. To generate the next documentary chapter based on the user's choices, the AI model is again invoked, taking their selections and profile information into account. Furthermore, the emotion engine uses the user's emotional data to tailor the next content and questions. This emotional data is processed in real time, ensuring the content best suited to the user's current emotional state is presented.
[0344] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0345] Device behavior
[0346] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0347] The device also has an emotion engine that analyzes the user's emotions in real time based on their facial expressions and voice. This emotion data is sent to the server and used to generate the next content. The newly generated documentary chapter is then sent back to the device and provided to the user.
[0348] User Actions
[0349] Users log in to the platform and watch the initial documentary video presented to them. They make their own choices in response to interactive questions displayed while watching. The emotion engine also analyzes the user's emotions, so emotional information is fed back to the system without any special operations. After watching, users provide feedback by entering their impressions and ratings into their device. This feedback is reflected in the next video generation.
[0350] Specific examples
[0351] Let's say the user is interested in "history." Based on this information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0352] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0353] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[0354] Hardware and software used
[0355] The server uses a cloud computing infrastructure with high-performance processors and large memory capacities. The generative AI model uses a deep learning model using the Python programming language and the TensorFlow library. The emotion engine combines facial expression analysis software and voice recognition technology. A NoSQL database is suitable for the database. The terminals are mobile devices such as smartphones, tablets, and smart glasses, which communicate with the server in real time via an internet connection.
[0356] Prompt Sentence Examples
[0357] Below are some example prompts to input to the generative AI model:
[0358] "User is interested in Ancient Egypt. Generate the next documentary chapter."
[0359] "User is interested in 'Pyramids'. Please generate the following details:"
[0360] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0361] Step 1:
[0362] The server retrieves the user's profile information from a database.
[0363] Input: User ID
[0364] Data processing and calculation: Issue database queries to extract user interests and past viewing history.
[0365] Output: User profile information
[0366] Step 2:
[0367] The server generates a prompt sentence for the generated AI model based on the profile information.
[0368] Input: User profile information
[0369] Data processing and calculation: Creates appropriate prompt statements in text format based on profile information.
[0370] Output: Prompt sentence to be passed to the generative AI model
[0371] Step 3:
[0372] The server inputs prompt sentences into the generative AI model to generate initial documentary content.
[0373] Input: prompt statement
[0374] Data processing and calculation: The generative AI model analyzes the prompt text and generates documentary content.
[0375] Output: Early documentary content
[0376] Step 4:
[0377] The server transmits the generated documentary content to the terminal.
[0378] Input: Early documentary content
[0379] Data processing and calculation: Content is optimized for the device and packetized.
[0380] Output: Data sent to the terminal
[0381] Step 5:
[0382] The terminal provides the received documentary content to the user and displays interactive questions.
[0383] Input: Early documentary content
[0384] Data processing and calculation: Play content and display interactive questions as an overlay.
[0385] Output: The content the user views and the questions
[0386] Step 6:
[0387] The user makes selections in response to interactive questions that are displayed.
[0388] Input: Interactive Question
[0389] Data processing and calculation: Obtain user selections via touch input or voice input.
[0390] Output: User's choice
[0391] Step 7:
[0392] The terminal transmits the user's selection to the server.
[0393] Input: User's choice
[0394] Data processing and calculation: The selected data is packetized for transmission to the server.
[0395] Output: Send selected data
[0396] Step 8:
[0397] The server receives the user's selection and generates the next prompt sentence along with the analysis results from the emotion engine.
[0398] Input: User choices and emotion data
[0399] Data processing and calculation: Based on the selection and emotional state, the following prompt sentence is created in text format.
[0400] Output: Next prompt statement
[0401] Step 9:
[0402] The server inputs the following prompt sentence into the generative AI model and generates the following documentary content.
[0403] Input: the following prompt statement
[0404] Data processing and calculation: The generative AI model analyzes the prompt text and generates new documentary content.
[0405] Output: Continuing documentary content
[0406] Step 10:
[0407] The server transmits the generated continuous documentary content to the terminal.
[0408] Input: Documentary content to be continued
[0409] Data processing and calculation: Content is optimized for the device and packetized.
[0410] Output: Data sent to the terminal
[0411] Step 11:
[0412] The device analyzes the user's emotional data in real time and sends the results back to the server.
[0413] Input: User's facial expression data and voice data
[0414] Data processing and calculation: Analyze using the emotion engine and obtain emotion data.
[0415] Output: Sending emotion data
[0416] Step 12:
[0417] The server stores the acquired emotional data and user feedback in a database and uses it to generate the next content.
[0418] Input: Emotion data and user feedback
[0419] Data processing and calculation: Index and store in a database.
[0420] Output: Data stored in the database
[0421] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0422] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0423] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0424] [Second embodiment]
[0425] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0426] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0427] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0428] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0429] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0431] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0432] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0433] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0434] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0435] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0436] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0437] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0438] Server Operation
[0439] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0440] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0441] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0442] Device behavior
[0443] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0444] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0445] User Actions
[0446] Users log in to the platform and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that appear, which determine the content of the next video. After watching, users provide feedback that will be reflected in the next video. Through this interactive selection and feedback process, users can experience infinitely customized documentary content.
[0447] Specific examples
[0448] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0449] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0450] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0451] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0452] The processing flow will be explained below.
[0453] Step 1:
[0454] A user logs into the platform. The user enters their authentication information on the login screen and is authenticated.
[0455] Step 2:
[0456] The terminal sends the user's login information to the server, which then retrieves the user's profile information from the database based on the received authentication information.
[0457] Step 3:
[0458] The server extracts the user's interests and past viewing history from their profile information and uses that information to send an initial documentary content generation request to the generative AI model.
[0459] Step 4:
[0460] The generative AI model generates an initial documentary video and returns it to the server, which then receives the video file and sends it to the device.
[0461] Step 5:
[0462] The terminal plays the initial documentary video and provides a viewing interface to the user, who then watches the video.
[0463] Step 6:
[0464] The server generates interactive questions based on the video content and sends the questions and options to the device, which then displays them to the user.
[0465] Step 7:
[0466] The user selects one of the options for the question presented, and the selection information is sent from the terminal to the server.
[0467] Step 8:
[0468] The server receives the user's selection and sends another request to the generative AI model to generate the next documentary chapter, updating the profile information if necessary.
[0469] Step 9:
[0470] The server receives the generated video file of the next documentary chapter and sends it to the device, which plays this new video.
[0471] Step 10:
[0472] The user watches a new documentary chapter and makes selections in response to the interactive questions that are presented again, and the process is repeated.
[0473] Step 11:
[0474] The server indexes and stores each generated video in a database so that it can be quickly provided to other users who make similar selections.
[0475] Step 12:
[0476] After watching a video, the user inputs feedback into the device, which then sends the feedback to the server and uses it to generate the next content.
[0477] This allows users to experience an infinite amount of customized documentary content.
[0478] Example 1
[0479] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0480] Conventional video content is difficult to customize according to user interests, and viewers often become bored with the same content. Furthermore, while systems existed that provided interactive experiences, they lacked the ability to dynamically generate content based on user choices. As a result, the appeal of the viewing experience was limited, and learning benefits were reduced.
[0481] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0482] In this invention, the server includes means for acquiring user profile information, means for generating interactive content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selections, means for indexing and storing each generated chapter in a database, and means for reusing the generated content to quickly provide it to other users who make the same selections, thereby enabling dynamic and infinite generation of customized documentary content according to the user's interests.
[0483] "User profile information" is data that indicates the individual attributes and behavioral patterns of each user, such as the user's interests, past viewing history, and selection history.
[0484] "Generative AI model" refers to an artificial intelligence algorithm or framework for dynamically generating documentary content based on user profile information and selection information.
[0485] An "interactive question" is a question that is displayed while the user is watching, prompting the user to make a choice, and is an element that influences the generation of the next content.
[0486] "Indexing" is the process of organizing and classifying generated content into a database so that it can be efficiently searched and referenced.
[0487] "Reuse" refers to reducing the computational cost of the system by making content once generated available to other users who make the same or similar choices.
[0488] "Feedback" refers to ratings and comments provided by users after viewing a content, and is sent to the server as a reference for the next content generation.
[0489] "Terminal" refers to an apparatus or device through which a user views content and inputs selections and feedback.
[0490] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0491] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history, which serves as the basic data for content generation. The server then generates documentary content using a generative AI model (for example, OpenAI's GPT-4 or TensorFlow). The prompt uses a format such as "Generate the content of the first interactive documentary based on the user's profile information." The generated content is then sent to the device.
[0492] The user uses a device to watch this initial documentary content. While watching, the server generates interactive questions, sends them to the device, and presents them to the user. For example, in response to the question, "Which part are you interested in next?", options such as "Pyramids," "Pharaoh's History," and "Daily Life" are presented. The user's selection is sent from the device to the server, and the next content is generated based on this. Again, the prompt is in the form, "Based on the user's selection and profile information, generate the next chapter of the documentary about the pyramids."
[0493] Each generated chapter is stored and indexed in a database on the server, allowing it to be quickly provided to other users who make the same selection, reducing computational costs. The content generated through reuse also corresponds to other users' selections.
[0494] After watching, users can enter their feedback into their device, which then sends it to the server. The server then incorporates this feedback into future content generation. For example, if a user gives a high rating to the "Pyramid" chapter, the server will take this into account in future content generation and adjust the content generation to provide more detailed information.
[0495] As a concrete example, if a user is interested in "history," the server generates a documentary on the theme of "ancient Egypt" based on the user's profile information. The prompt text used is, "Based on the user's profile information, generate an interactive documentary about ancient Egypt." The generated video is displayed on the device, and the user begins watching. Interactive questions are presented during viewing, and the next documentary chapter is generated based on the user's selection. In this way, an infinite documentary experience completely customized to the user's interests is realized.
[0496] This system allows users to enjoy an interactive and dynamic viewing experience by providing documentary content customized to their interests, thereby preventing user boredom and maximizing learning effectiveness.
[0497] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0498] Step 1:
[0499] The server receives the user's identification information (user ID) when the user logs into the platform. As input, the user ID is sent to the server. The server uses this user ID to retrieve the corresponding profile information (interests, viewing history, etc.) from the database. As output, the retrieved profile information is ready to be passed to the generative AI model.
[0500] Step 2:
[0501] The server uses the acquired profile information to create a prompt appropriate for the generative AI model. This prompt includes instructions for generating initial documentary content based on the user's interests. For example, a prompt might be created such as, "Generate an interactive documentary about ancient Egypt based on the user's profile information." The acquired profile information is used as input, and the prompt is generated as output.
[0502] Step 3:
[0503] The server inputs the generated prompt sentences into a generative AI model to generate initial documentary content. The generative AI model generates data such as text, images, and videos based on the input prompt sentences. The output is the initial documentary content, which includes content that is in line with the user's interests.
[0504] Step 4:
[0505] The server sends the generated initial documentary content to the terminal, which uses the generated documentary content as input and transfers it to the terminal as output, in an appropriate format (e.g., video file or streaming data).
[0506] Step 5:
[0507] The terminal displays the initial documentary content received from the server to the user, who then begins viewing the content, using the documentary content received from the server as input and providing viewing to the user as output.
[0508] Step 6:
[0509] The server generates interactive questions according to the viewing progress. For example, when a specific scene in a video is reached, the server generates a question such as "Which part are you interested in next?" and presents options such as "Pyramids," "Pharaoh's history," and "Daily life." The viewing progress data is used as input, and the interactive questions are generated as output.
[0510] Step 7:
[0511] The terminal presents the user with interactive questions sent from the server. The user selects one of the options presented. The user's option is entered into the terminal as input, and the selection is sent to the server as output.
[0512] Step 8:
[0513] The server sends a prompt to the generative AI model again based on the user's selection and profile information to generate the next chapter of documentary content. The prompt might be something like, "Generate the next chapter of a documentary about pyramids based on the user's selection and profile information." The user's selection and profile information are used as input, and the documentary content for the next chapter is generated as output.
[0514] Step 9:
[0515] The server sends the generated documentary content of the next chapter to the terminal. The terminal receives this new content and provides it to the user. The generated documentary content is used as input and sent to the terminal as output.
[0516] Step 10:
[0517] A user watches new documentary content and then provides feedback. As input, the feedback information is input to the terminal, and as output, the feedback is sent to the server.
[0518] Step 11:
[0519] The server stores the feedback received from the device in a database and uses it for the next content generation. The user's feedback information is used as input and stored in the database as output. This feedback information is used to adjust the prompt text to be generated next time.
[0520] Step 12:
[0521] The server analyzes the accumulated feedback and reflects it in the next content generation. This allows for more highly customized content based on the user's interests and satisfaction. The feedback data is used as input, and the output is used to adjust the generated prompts and the generative AI model.
[0522] (Application example 1)
[0523] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0524] Conventional content delivery systems have difficulty generating and providing personalized content based on user interests and preferences, resulting in inconsistent viewing experiences. Furthermore, the generated content lacks interactivity, making it difficult to maintain user interest over the long term. Furthermore, there is a lack of a way to utilize user feedback to improve the quality of the viewing experience.
[0525] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0526] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, and means for delivering the generated content via a smartphone application and collecting user feedback, thereby enabling the generation of personalized content based on the user's interests, providing an interactive experience, and utilizing the collected feedback for future content generation.
[0527] "User profile information" is information including the user's interests, past viewing history, personal data, etc., and is the basis for content generation.
[0528] A "generative AI model" is a model that uses artificial intelligence to generate content based on user profile information and interactive choices.
[0529] "Documentary content" refers to content that is comprised of fact-based video and audio, and includes personalized content tailored to the user's interests.
[0530] An "interactive question" is a choice or question presented to generate different content based on a user's selection.
[0531] A "smartphone application" is an application that runs on a smartphone and provides an interface for users to answer interactive questions and view content.
[0532] "Feedback" refers to the act of a user providing opinions and impressions about content they have viewed, and the information provided can be used to help generate the next piece of content.
[0533] "Indexing" is the process of organizing generated content and storing it in a database for easy searching and reuse.
[0534] A system for implementing this invention acquires user profile information and uses that information to generate interactive documentary content based on the user's interests and preferences using a generative AI model. The system is composed of the following components:
[0535] 1. Server Operation
[0536] The server retrieves the user's profile information from a database, including their interests and past viewing history. Based on this profile information, the server provides appropriate input data to a generative AI model to generate the initial documentary content. The generative AI model used is OpenAI GPT-3, a general artificial intelligence platform. The generated content is then sent to the device, where the user can begin watching.
[0537] 2. Device Operation
[0538] The user's smartphone application receives the initial documentary content from the server. When the user begins watching, the application displays interactive questions. The user's selections in response to these questions are sent from the device to the server. To generate the next documentary chapter, the server again invokes the generative AI model, inputting the user's selections and profile information. The newly generated documentary chapter is again sent to the device and served to the user.
[0539] 3. User Actions
[0540] Users log in to the smartphone application and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that determine the content of the next video. After watching, users provide feedback, and that information is reflected in the next content generation.
[0541] Specific examples
[0542] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0543] While watching, the question "Which part are you interested in next?" is displayed, and the options are "Pyramids," "Pharaoh's History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focusing on "Pyramids." Next, the following sentence is used as an example of a prompt sentence for the generative AI model:
[0544] Example prompt sentence:
[0545] "Username is interested in history, particularly ancient Egypt. Please generate a 10-minute documentary on the subject."
[0546] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0547] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0548] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0549] Step 1:
[0550] The server retrieves user profile information from a database. The input is a user ID or other identifying information, and the output is profile information including the user's interests and past viewing history. It filters and retrieves the required data from the database (e.g., MySQL) using SQL queries.
[0551] Step 2:
[0552] The server generates a content generation prompt for the generative AI model based on the acquired profile information. The input is the user profile information, and the output is a prompt sentence for the generative AI model. Specifically, the server analyzes the profile information using a Python script and constructs an appropriate prompt sentence.
[0553] Step 3:
[0554] The server inputs a prompt into a generative AI model (e.g., OpenAI GPT-3) to generate the initial documentary content. The input is the prompt, and the output is the generated documentary content. Specifically, the server makes an API call and receives the generated text and video data.
[0555] Step 4:
[0556] The server sends the generated content to the user's smartphone app. The input is the generated documentary content, and the output is data transmission to the smartphone app. Data is sent using an HTTP request.
[0557] Step 5:
[0558] The device (smartphone app) displays the received documentary content to the user. The user begins watching and is presented with interactive questions. The input is the content received from the server, and the output is the interface displayed to the user. Specifically, the display uses UI elements created with React Native or Flutter.
[0559] Step 6:
[0560] The user answers the interactive questions presented to them. The input is the interactive question, and the output is the user's answer. When the user presses the selection button, the selection is registered within the application.
[0561] Step 7:
[0562] The terminal sends the user's answer information to the server. The input is the user's answer information, and the output is data sent to the server. The data is sent using an HTTP POST request.
[0563] Step 8:
[0564] The server uses the user's selections to generate the next documentary chapter. The input is the user's selections and profile information, and the output is a new prompt for the generative AI model and new generated content. Another API call is made to retrieve the necessary data.
[0565] Step 9:
[0566] The generated new content is then sent to the device again, and the user begins viewing it. The input is the generated new content, and the output is data sent to the smartphone app.
[0567] Step 10:
[0568] The user views new content and answers the interactive questions again. This process is repeated between the server, the device, and the user. The input is the new interactive questions, and the output is the user's new answers.
[0569] This allows users to enjoy an endless variety of personalized documentary experiences.
[0570] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0571] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[0572] Server Operation
[0573] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0574] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0575] Furthermore, the emotion engine uses the user's emotional data to tailor the next content or question. This emotional data is processed in real time, so the content provided is most appropriate for the user's current emotional state.
[0576] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0577] Device behavior
[0578] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0579] The device is equipped with an emotion engine that analyzes the user's emotions in real time from their facial expressions, voice, etc. This emotion data is sent to the server and used to generate the next content.
[0580] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0581] User Actions
[0582] Users log in to the platform and watch the initial documentary video provided. They make their own choices in response to interactive questions that appear while watching. The emotion engine also analyzes the user's emotions, and the emotional information is fed back to the system without any special operation.
[0583] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[0584] Specific examples
[0585] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0586] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0587] Additionally, the emotional state of the user while watching is analyzed by the emotion engine, and if the user is excited, for example, more detailed and visually stimulating content is provided.
[0588] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to their interests and emotions.
[0589] The system customizes documentaries according to individual users' interests and emotional state, providing an interactive viewing experience that combats viewer boredom and maximizes learning outcomes.
[0590] The processing flow will be explained below.
[0591] Server Processing
[0592] Step 1:
[0593] The server receives the user's login request and verifies the credentials.
[0594] Step 2:
[0595] The server retrieves the user's profile information from a database, which includes the user's interests, past viewing history, and individual viewing habits.
[0596] Step 3:
[0597] The server provides profile information to the generative AI model to generate the initial documentary content, which then generates the initial video based on this information.
[0598] Step 4:
[0599] The server receives the generated initial documentary video and transmits it to the terminal.
[0600] Step 5:
[0601] The server also receives data from the emotion engine to capture the user's current emotional state during video playback. The emotion data is processed in real time to influence the next content generation.
[0602] Terminal handling
[0603] Step 1:
[0604] The terminal displays an interface for the user to enter login information, which is then sent to the server.
[0605] Step 2:
[0606] The terminal plays back the initial documentary video received from the server and provides it to the user.
[0607] Step 3:
[0608] The emotion engine built into the device analyzes the user's emotions in real time from their facial expressions and voice, and this emotion data is sent to the server.
[0609] Step 4:
[0610] Interactive questions are displayed during the video, which contain multiple options for the user to select.
[0611] Step 5:
[0612] The user's selection results and emotion data are sent to the server and used as data for generating new content.
[0613] User Action
[0614] Step 1:
[0615] Users log in to the platform and watch the initial video.
[0616] Step 2:
[0617] While watching, interactive questions are displayed and the user is prompted to select one of the options.
[0618] Step 3:
[0619] The emotion engine analyzes the user's emotions in real time and sends the data to the server.
[0620] Step 4:
[0621] The user then watches the next newly generated video and answers the questions again.
[0622] Step 5:
[0623] After viewing, the user enters feedback, and this data is also sent to the server.
[0624] Specific examples
[0625] Step 1:
[0626] When a user logs in with an interest in "history," the server obtains this interest information.
[0627] Step 2:
[0628] The server requests an initial documentary related to "Ancient Egypt" from the generative AI model, while the device detects the user's emotional state, analyzing facial expressions and voice sounds that indicate excitement or curiosity in real time.
[0629] Step 3:
[0630] An initial video is generated and provided to the user. While watching the video, the user is prompted with the question, "Which part are you interested in next?" The options offered are "Pyramids," "Pharaoh's History," and "Everyday Life."
[0631] Step 4:
[0632] When the user selects "Pyramid" and the emotion engine detects the user's state of excitement, the server adjusts the next video it generates to be more visually stimulating.
[0633] Step 5:
[0634] A new video is generated and sent to the user's device, the user starts watching again, and the same process continues.
[0635] This system allows users to experience a wide variety of documentary content optimized for their emotional state.
[0636] Example 2
[0637] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0638] While conventional documentary content generation systems have the ability to generate content based on user interests and selection, they do not fully utilize user emotions or real-time feedback. This makes it difficult to fine-tune personalization based on individual user interests and emotional states, resulting in a poor viewing experience. Furthermore, it is difficult to efficiently reuse generated content, resulting in a waste of computing resources.
[0639] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0640] In this invention, the server includes a means for acquiring user profile information, a means for generating documentary content based on the user's interests and selections using a generative AI model, a means for collecting emotion data in real time using an emotion engine that recognizes the user's emotions and using that data to generate subsequent content, and a means for indexing and storing the generated content in a database. This makes it possible to provide highly personalized documentary content based on the user's interests and emotions, thereby realizing efficient use of computing resources and an improved quality of viewing experience.
[0641] "User profile information" is information that includes personal data such as a user's interests, past viewing history, preferences, and the like.
[0642] A "generative AI model" is an artificial intelligence system that generates appropriate documentary content based on a user's profile information and choices.
[0643] "Documentary content" is content consisting of video, audio and text based on facts and information.
[0644] An "interactive question" is a question with options that is presented to the user while they are watching, and the user's choice is reflected in the next content.
[0645] The "emotion engine" is a technology that collects and analyzes emotional data from users' facial expressions, voice, etc. in real time.
[0646] "Emotion data" is data that indicates the user's current emotional state, and is used to generate the next content.
[0647] A "database" is a system that efficiently stores and manages acquired information and generated content.
[0648] "Indexing" is the process of organizing data into a searchable format that makes queries within a database more efficient.
[0649] "Feedback" refers to information such as impressions and evaluations provided by users after viewing a content, and is useful for creating the next content.
[0650] "Reuse" refers to the reuse of content that has been generated once in a manner that corresponds to the selections of other users.
[0651] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[0652] Server Operation
[0653] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history. The server then provides this information to a generative AI model to generate the initial documentary content. The generative AI model can be OpenAI GPT-4 or similar.
[0654] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The generative AI model is again invoked to generate the next documentary chapter based on the user's choices, inputting their selections and profile information.
[0655] Furthermore, the system uses the user's emotional data recognized by the emotion engine to tailor the next content or question. Based on the emotional data processed in real time, the system provides content that best suits the user's current emotional state. Each generated chapter is stored and indexed in a database, allowing it to be quickly provided to other users who make the same selection, reducing computational costs.
[0656] Device behavior
[0657] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0658] The device also has a built-in emotion engine that analyzes emotions from the user's facial expressions and voice in real time and sends the data to the server. This emotion data is used to generate new content. The newly generated documentary chapter is sent to the device and provided to the user. Feedback entered by the user after watching the video is also sent from the device to the server and used to generate the next piece of content.
[0659] User Actions
[0660] Users log in to the platform and watch the initial documentary video provided to them. They then select options in response to interactive questions that appear while they are watching. Furthermore, emotional information analyzed by the emotion engine is fed back to the system in real time, so no special operations are required.
[0661] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[0662] Specific examples
[0663] For example, if a user is interested in "history," the server uses this interest information to request the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary. While watching, the question "Which part are you interested in next?" is displayed, and the options are presented: "Pyramids," "Pharaoh's history," and "Daily life." If the user selects "Pyramids," the selection is sent to the server, and the next chapter of the documentary is generated with content focused on "Pyramids."
[0664] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[0665] Prompt Sentence Examples
[0666] "If a user is interested in history, we can have the generative AI model generate a documentary on the theme of 'Ancient Egypt,' taking into account the user's specific profile information and past viewing history to provide interesting content."
[0667] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0668] Step 1:
[0669] A user logs in to the platform. The user accesses the platform and enters their username and password on the login screen. When they press the login button, the information is sent to the server. The server performs authentication, and if successful, starts the user's session.
[0670] Input: Username, Password
[0671] Output: User authentication result (success / failure), session start
[0672] Step 2:
[0673] The server retrieves the user's profile information: After successful authentication, the server retrieves the user's profile information (interests, past viewing history, etc.) from the database. This profile information is used in the next processing step.
[0674] Input: User ID
[0675] Output: User profile information
[0676] Step 3:
[0677] The server provides prompts to the generative AI model to generate initial documentary content. The server creates appropriate prompt sentences for the generative AI model based on the profile information and sends them to the model. The generative AI model generates documentary content based on these prompt sentences.
[0678] Input: User profile information, prompt text
[0679] Output: Initial documentary content
[0680] Step 4:
[0681] The server sends the generated documentary content to the device. The server receives the content generated by the generative AI model and sends it to the user's device. The content is ready for the user to watch.
[0682] Input: Early documentary content
[0683] Output: Content delivery to devices
[0684] Step 5:
[0685] The terminal displays the documentary content and the user starts viewing it. The terminal plays the content received from the server and makes it available for viewing by the user. The user starts viewing the content.
[0686] Input: Documentary content from the server
[0687] Output: Playback begins, user watches
[0688] Step 6:
[0689] The terminal displays interactive questions. While the user is watching the documentary, the terminal displays the interactive questions received from the server and presents the user with options.
[0690] Input: Interactive questions from the server
[0691] Output: Screen presenting options to the user
[0692] Step 7:
[0693] The user selects an option for an interactive question. The user selects one option from multiple options and enters it into the terminal. The terminal sends the selected information to the server.
[0694] Input: User's choice
[0695] Output: Sending choices from the terminal to the server
[0696] Step 8:
[0697] The terminal transmits the selection to the server. After receiving the user's selection, the terminal transmits the data to the server, providing input data for the next content generation.
[0698] Input: User's choice
[0699] Output: Send selected data to the server
[0700] Step 9:
[0701] The server requests a new documentary chapter from the generative AI model. The server generates a prompt for the new documentary chapter based on the user's choices and sends it to the generative AI model. The generative AI model then generates the new chapter based on that prompt.
[0702] Input: User choices, prompt text
[0703] Output: New documentary chapter
[0704] Step 10:
[0705] The device collects the user's emotional data and sends it to the server. The device's built-in emotion engine analyzes the user's emotional data while watching and sends the information to the server.
[0706] Input: User's emotional state
[0707] Output: Send emotion data to the server
[0708] Step 11:
[0709] The server generates the next content based on the emotion data. It analyzes the received emotion data and adjusts the content and questions of the next documentary chapter. The generated documentary chapter is indexed and stored in a database.
[0710] Input: User emotion data
[0711] Output: adjusted documentary chapters, stored in a database
[0712] Step 12:
[0713] The server sends new documentary chapters to the terminal, and the server sends the new chapters to the terminal so that the user can continue watching.
[0714] Enter: a new documentary chapter
[0715] Output: New documentary streaming to your device
[0716] Step 13:
[0717] User continues watching. The user watches the newly received documentary chapter and continues their viewing experience.
[0718] Enter: a new documentary chapter
[0719] Output: User continues watching
[0720] Step 14:
[0721] Users provide feedback on videos. After watching, users input their impressions and ratings into their devices, and the feedback is sent to the server. This information is used to generate new content for the next time.
[0722] Input: User feedback
[0723] Output: Send feedback to the server
[0724] The above processing steps enable the system of the present invention to provide highly personalized documentary content according to the user's interests, preferences, and emotional state.
[0725] (Application example 2)
[0726] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0727] Conventional documentary content generation systems personalize content based on user interests and selections, but they fail to take into account the user's emotional state. This makes it difficult to provide content that adapts to the user's emotional state, and it is therefore difficult to significantly improve the satisfaction of the viewing experience. Furthermore, user feedback is often not adequately reflected in the next content generation, limiting the improvement of content quality.
[0728] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0729] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selection, means for indexing and storing the generated content in a database, and means for recognizing the user's emotions using an emotion engine and reflecting the emotion data in the generation of the next content, thereby enabling the provision of sophisticated personalized content according to the user's emotional state.
[0730] "User profile information" is individual information about a user, such as the user's interests and past viewing history.
[0731] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate documentary content based on user interests and choices.
[0732] An "interactive question" is a question that is presented to a user to elicit a choice from the user.
[0733] The "emotion engine" is a system that analyzes the user's facial expressions and voice to obtain emotional data.
[0734] "Documentary content" refers to fact-based video works and informational content.
[0735] "Feedback" refers to data such as reactions, evaluations, and opinions from users.
[0736] "Personalization" is a method of providing individualized experiences and content.
[0737] A "database" is a system in which information is systematically stored and can be retrieved as needed.
[0738] "Indexing" is the process of organizing information in a database so that it can be quickly searched.
[0739] The system of the present invention acquires user profile information and generates interactive documentary content using a generative AI model based on that information. By combining it with an emotion engine, it is possible to provide more sophisticated personalized content.
[0740] Server Operation
[0741] The server first retrieves the user's profile information from a database, including the user's interests and past viewing history. The server then uses this profile information to provide appropriate input data to a generative AI model to generate initial documentary content. The generated documentary content is then sent to the device, where the user can begin viewing.
[0742] Interactive questions are generated and presented to the user as they watch. To generate the next documentary chapter based on the user's choices, the AI model is again invoked, taking their selections and profile information into account. Furthermore, the emotion engine uses the user's emotional data to tailor the next content and questions. This emotional data is processed in real time, ensuring the content best suited to the user's current emotional state is presented.
[0743] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0744] Device behavior
[0745] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0746] The device also has an emotion engine that analyzes the user's emotions in real time based on their facial expressions and voice. This emotion data is sent to the server and used to generate the next content. The newly generated documentary chapter is then sent back to the device and provided to the user.
[0747] User Actions
[0748] Users log in to the platform and watch the initial documentary video presented to them. They make their own choices in response to interactive questions displayed while watching. The emotion engine also analyzes the user's emotions, so emotional information is fed back to the system without any special operations. After watching, users provide feedback by entering their impressions and ratings into their device. This feedback is reflected in the next video generation.
[0749] Specific examples
[0750] Let's say the user is interested in "history." Based on this information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0751] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0752] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[0753] Hardware and software used
[0754] The server uses a cloud computing infrastructure with high-performance processors and large memory capacities. The generative AI model uses a deep learning model using the Python programming language and the TensorFlow library. The emotion engine combines facial expression analysis software and voice recognition technology. A NoSQL database is suitable for the database. The terminals are mobile devices such as smartphones, tablets, and smart glasses, which communicate with the server in real time via an internet connection.
[0755] Prompt Sentence Examples
[0756] Below are some example prompts to input to the generative AI model:
[0757] "User is interested in Ancient Egypt. Generate the next documentary chapter."
[0758] "User is interested in 'Pyramids'. Please generate the following details:"
[0759] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0760] Step 1:
[0761] The server retrieves the user's profile information from a database.
[0762] Input: User ID
[0763] Data processing and calculation: Issue database queries to extract user interests and past viewing history.
[0764] Output: User profile information
[0765] Step 2:
[0766] The server generates a prompt sentence for the generated AI model based on the profile information.
[0767] Input: User profile information
[0768] Data processing and calculation: Creates appropriate prompt statements in text format based on profile information.
[0769] Output: Prompt sentence to be passed to the generative AI model
[0770] Step 3:
[0771] The server inputs prompt sentences into the generative AI model to generate initial documentary content.
[0772] Input: prompt statement
[0773] Data processing and calculation: The generative AI model analyzes the prompt text and generates documentary content.
[0774] Output: Early documentary content
[0775] Step 4:
[0776] The server transmits the generated documentary content to the terminal.
[0777] Input: Early documentary content
[0778] Data processing and calculation: Content is optimized for the device and packetized.
[0779] Output: Data sent to the terminal
[0780] Step 5:
[0781] The terminal provides the received documentary content to the user and displays interactive questions.
[0782] Input: Early documentary content
[0783] Data processing and calculation: Play content and display interactive questions as an overlay.
[0784] Output: The content the user views and the questions
[0785] Step 6:
[0786] The user makes selections in response to interactive questions that are displayed.
[0787] Input: Interactive Question
[0788] Data processing and calculation: Obtain user selections via touch input or voice input.
[0789] Output: User's choice
[0790] Step 7:
[0791] The terminal transmits the user's selection to the server.
[0792] Input: User's choice
[0793] Data processing and calculation: The selected data is packetized for transmission to the server.
[0794] Output: Send selected data
[0795] Step 8:
[0796] The server receives the user's selection and generates the next prompt sentence along with the analysis results from the emotion engine.
[0797] Input: User choices and emotion data
[0798] Data processing and calculation: Based on the selection and emotional state, the following prompt sentence is created in text format.
[0799] Output: Next prompt statement
[0800] Step 9:
[0801] The server inputs the following prompt sentence into the generative AI model and generates the following documentary content.
[0802] Input: the following prompt statement
[0803] Data processing and calculation: The generative AI model analyzes the prompt text and generates new documentary content.
[0804] Output: Continuing documentary content
[0805] Step 10:
[0806] The server transmits the generated continuous documentary content to the terminal.
[0807] Input: Documentary content to be continued
[0808] Data processing and calculation: Content is optimized for the device and packetized.
[0809] Output: Data sent to the terminal
[0810] Step 11:
[0811] The device analyzes the user's emotional data in real time and sends the results back to the server.
[0812] Input: User's facial expression data and voice data
[0813] Data processing and calculation: Analyze using the emotion engine and obtain emotion data.
[0814] Output: Sending emotion data
[0815] Step 12:
[0816] The server stores the acquired emotional data and user feedback in a database and uses it to generate the next content.
[0817] Input: Emotion data and user feedback
[0818] Data processing and calculation: Index and store in a database.
[0819] Output: Data stored in the database
[0820] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0821] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0822] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0823] [Third embodiment]
[0824] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0825] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0826] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0827] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0828] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0829] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0830] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0831] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0832] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0833] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0834] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0835] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0836] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0837] Server Operation
[0838] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0839] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0840] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0841] Device behavior
[0842] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0843] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0844] User Actions
[0845] Users log in to the platform and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that appear, which determine the content of the next video. After watching, users provide feedback that will be reflected in the next video. Through this interactive selection and feedback process, users can experience infinitely customized documentary content.
[0846] Specific examples
[0847] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0848] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0849] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0850] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0851] The processing flow will be explained below.
[0852] Step 1:
[0853] A user logs into the platform. The user enters their authentication information on the login screen and is authenticated.
[0854] Step 2:
[0855] The terminal sends the user's login information to the server, which then retrieves the user's profile information from the database based on the received authentication information.
[0856] Step 3:
[0857] The server extracts the user's interests and past viewing history from their profile information and uses that information to send an initial documentary content generation request to the generative AI model.
[0858] Step 4:
[0859] The generative AI model generates an initial documentary video and returns it to the server, which then receives the video file and sends it to the device.
[0860] Step 5:
[0861] The terminal plays the initial documentary video and provides a viewing interface to the user, who then watches the video.
[0862] Step 6:
[0863] The server generates interactive questions based on the video content and sends the questions and options to the device, which then displays them to the user.
[0864] Step 7:
[0865] The user selects one of the options for the question presented, and the selection information is sent from the terminal to the server.
[0866] Step 8:
[0867] The server receives the user's selection and sends another request to the generative AI model to generate the next documentary chapter, updating the profile information if necessary.
[0868] Step 9:
[0869] The server receives the generated video file of the next documentary chapter and sends it to the device, which plays this new video.
[0870] Step 10:
[0871] The user watches a new documentary chapter and makes selections in response to the interactive questions that are presented again, and the process is repeated.
[0872] Step 11:
[0873] The server indexes and stores each generated video in a database so that it can be quickly provided to other users who make similar selections.
[0874] Step 12:
[0875] After watching a video, the user inputs feedback into the device, which then sends the feedback to the server and uses it to generate the next content.
[0876] This allows users to experience an infinite amount of customized documentary content.
[0877] Example 1
[0878] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0879] Conventional video content is difficult to customize according to user interests, and viewers often become bored with the same content. Furthermore, while systems existed that provided interactive experiences, they lacked the ability to dynamically generate content based on user choices. As a result, the appeal of the viewing experience was limited, and learning benefits were reduced.
[0880] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0881] In this invention, the server includes means for acquiring user profile information, means for generating interactive content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selections, means for indexing and storing each generated chapter in a database, and means for reusing the generated content to quickly provide it to other users who make the same selections, thereby enabling dynamic and infinite generation of customized documentary content according to the user's interests.
[0882] "User profile information" is data that indicates the individual attributes and behavioral patterns of each user, such as the user's interests, past viewing history, and selection history.
[0883] "Generative AI model" refers to an artificial intelligence algorithm or framework for dynamically generating documentary content based on user profile information and selection information.
[0884] An "interactive question" is a question that is displayed while the user is watching, prompting the user to make a choice, and is an element that influences the generation of the next content.
[0885] "Indexing" is the process of organizing and classifying generated content into a database so that it can be efficiently searched and referenced.
[0886] "Reuse" refers to reducing the computational cost of the system by making content once generated available to other users who make the same or similar choices.
[0887] "Feedback" refers to ratings and comments provided by users after viewing a content, and is sent to the server as a reference for the next content generation.
[0888] "Terminal" refers to an apparatus or device through which a user views content and inputs selections and feedback.
[0889] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[0890] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history, which serves as the basic data for content generation. The server then generates documentary content using a generative AI model (for example, OpenAI's GPT-4 or TensorFlow). The prompt uses a format such as "Generate the content of the first interactive documentary based on the user's profile information." The generated content is then sent to the device.
[0891] The user uses a device to watch this initial documentary content. While watching, the server generates interactive questions, sends them to the device, and presents them to the user. For example, in response to the question, "Which part are you interested in next?", options such as "Pyramids," "Pharaoh's History," and "Daily Life" are presented. The user's selection is sent from the device to the server, and the next content is generated based on this. Again, the prompt is in the form, "Based on the user's selection and profile information, generate the next chapter of the documentary about the pyramids."
[0892] Each generated chapter is stored and indexed in a database on the server, allowing it to be quickly provided to other users who make the same selection, reducing computational costs. The content generated through reuse also corresponds to other users' selections.
[0893] After watching, users can enter their feedback into their device, which then sends it to the server. The server then incorporates this feedback into future content generation. For example, if a user gives a high rating to the "Pyramid" chapter, the server will take this into account in future content generation and adjust the content generation to provide more detailed information.
[0894] As a concrete example, if a user is interested in "history," the server generates a documentary on the theme of "ancient Egypt" based on the user's profile information. The prompt text used is, "Based on the user's profile information, generate an interactive documentary about ancient Egypt." The generated video is displayed on the device, and the user begins watching. Interactive questions are presented during viewing, and the next documentary chapter is generated based on the user's selection. In this way, an infinite documentary experience completely customized to the user's interests is realized.
[0895] This system allows users to enjoy an interactive and dynamic viewing experience by providing documentary content customized to their interests, thereby preventing user boredom and maximizing learning effectiveness.
[0896] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0897] Step 1:
[0898] The server receives the user's identification information (user ID) when the user logs into the platform. As input, the user ID is sent to the server. The server uses this user ID to retrieve the corresponding profile information (interests, viewing history, etc.) from the database. As output, the retrieved profile information is ready to be passed to the generative AI model.
[0899] Step 2:
[0900] The server uses the acquired profile information to create a prompt appropriate for the generative AI model. This prompt includes instructions for generating initial documentary content based on the user's interests. For example, a prompt might be created such as, "Generate an interactive documentary about ancient Egypt based on the user's profile information." The acquired profile information is used as input, and the prompt is generated as output.
[0901] Step 3:
[0902] The server inputs the generated prompt sentences into a generative AI model to generate initial documentary content. The generative AI model generates data such as text, images, and videos based on the input prompt sentences. The output is the initial documentary content, which includes content that is in line with the user's interests.
[0903] Step 4:
[0904] The server sends the generated initial documentary content to the terminal, which uses the generated documentary content as input and transfers it to the terminal as output, in an appropriate format (e.g., video file or streaming data).
[0905] Step 5:
[0906] The terminal displays the initial documentary content received from the server to the user, who then begins viewing the content, using the documentary content received from the server as input and providing viewing to the user as output.
[0907] Step 6:
[0908] The server generates interactive questions according to the viewing progress. For example, when a specific scene in a video is reached, the server generates a question such as "Which part are you interested in next?" and presents options such as "Pyramids," "Pharaoh's history," and "Daily life." The viewing progress data is used as input, and the interactive questions are generated as output.
[0909] Step 7:
[0910] The terminal presents the user with interactive questions sent from the server. The user selects one of the options presented. The user's option is entered into the terminal as input, and the selection is sent to the server as output.
[0911] Step 8:
[0912] The server sends a prompt to the generative AI model again based on the user's selection and profile information to generate the next chapter of documentary content. The prompt might be something like, "Generate the next chapter of a documentary about pyramids based on the user's selection and profile information." The user's selection and profile information are used as input, and the documentary content for the next chapter is generated as output.
[0913] Step 9:
[0914] The server sends the generated documentary content of the next chapter to the terminal. The terminal receives this new content and provides it to the user. The generated documentary content is used as input and sent to the terminal as output.
[0915] Step 10:
[0916] A user watches new documentary content and then provides feedback. As input, the feedback information is input to the terminal, and as output, the feedback is sent to the server.
[0917] Step 11:
[0918] The server stores the feedback received from the device in a database and uses it for the next content generation. The user's feedback information is used as input and stored in the database as output. This feedback information is used to adjust the prompt text to be generated next time.
[0919] Step 12:
[0920] The server analyzes the accumulated feedback and reflects it in the next content generation. This allows for more highly customized content based on the user's interests and satisfaction. The feedback data is used as input, and the output is used to adjust the generated prompts and the generative AI model.
[0921] (Application example 1)
[0922] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0923] Conventional content delivery systems have difficulty generating and providing personalized content based on user interests and preferences, resulting in inconsistent viewing experiences. Furthermore, the generated content lacks interactivity, making it difficult to maintain user interest over the long term. Furthermore, there is a lack of a way to utilize user feedback to improve the quality of the viewing experience.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0925] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, and means for delivering the generated content via a smartphone application and collecting user feedback, thereby enabling the generation of personalized content based on the user's interests, providing an interactive experience, and utilizing the collected feedback for future content generation.
[0926] "User profile information" is information including the user's interests, past viewing history, personal data, etc., and is the basis for content generation.
[0927] A "generative AI model" is a model that uses artificial intelligence to generate content based on user profile information and interactive choices.
[0928] "Documentary content" refers to content that is comprised of fact-based video and audio, and includes personalized content tailored to the user's interests.
[0929] An "interactive question" is a choice or question presented to generate different content based on a user's selection.
[0930] A "smartphone application" is an application that runs on a smartphone and provides an interface for users to answer interactive questions and view content.
[0931] "Feedback" refers to the act of a user providing opinions and impressions about content they have viewed, and the information provided can be used to help generate the next piece of content.
[0932] "Indexing" is the process of organizing generated content and storing it in a database for easy searching and reuse.
[0933] A system for implementing this invention acquires user profile information and uses that information to generate interactive documentary content based on the user's interests and preferences using a generative AI model. The system is composed of the following components:
[0934] 1. Server Operation
[0935] The server retrieves the user's profile information from a database, including their interests and past viewing history. Based on this profile information, the server provides appropriate input data to a generative AI model to generate the initial documentary content. The generative AI model used is OpenAI GPT-3, a general artificial intelligence platform. The generated content is then sent to the device, where the user can begin watching.
[0936] 2. Device Operation
[0937] The user's smartphone application receives the initial documentary content from the server. When the user begins watching, the application displays interactive questions. The user's selections in response to these questions are sent from the device to the server. To generate the next documentary chapter, the server again invokes the generative AI model, inputting the user's selections and profile information. The newly generated documentary chapter is again sent to the device and served to the user.
[0938] 3. User Actions
[0939] Users log in to the smartphone application and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that determine the content of the next video. After watching, users provide feedback, and that information is reflected in the next content generation.
[0940] Specific examples
[0941] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0942] While watching, the question "Which part are you interested in next?" is displayed, and the options are "Pyramids," "Pharaoh's History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focusing on "Pyramids." Next, the following sentence is used as an example of a prompt sentence for the generative AI model:
[0943] Example prompt sentence:
[0944] "Username is interested in history, particularly ancient Egypt. Please generate a 10-minute documentary on the subject."
[0945] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[0946] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[0947] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0948] Step 1:
[0949] The server retrieves user profile information from a database. The input is a user ID or other identifying information, and the output is profile information including the user's interests and past viewing history. It filters and retrieves the required data from the database (e.g., MySQL) using SQL queries.
[0950] Step 2:
[0951] The server generates a content generation prompt for the generative AI model based on the acquired profile information. The input is the user profile information, and the output is a prompt sentence for the generative AI model. Specifically, the server analyzes the profile information using a Python script and constructs an appropriate prompt sentence.
[0952] Step 3:
[0953] The server inputs a prompt into a generative AI model (e.g., OpenAI GPT-3) to generate the initial documentary content. The input is the prompt, and the output is the generated documentary content. Specifically, the server makes an API call and receives the generated text and video data.
[0954] Step 4:
[0955] The server sends the generated content to the user's smartphone app. The input is the generated documentary content, and the output is data transmission to the smartphone app. Data is sent using an HTTP request.
[0956] Step 5:
[0957] The device (smartphone app) displays the received documentary content to the user. The user begins watching and is presented with interactive questions. The input is the content received from the server, and the output is the interface displayed to the user. Specifically, the display uses UI elements created with React Native or Flutter.
[0958] Step 6:
[0959] The user answers the interactive questions presented to them. The input is the interactive question, and the output is the user's answer. When the user presses the selection button, the selection is registered within the application.
[0960] Step 7:
[0961] The terminal sends the user's answer information to the server. The input is the user's answer information, and the output is data sent to the server. The data is sent using an HTTP POST request.
[0962] Step 8:
[0963] The server uses the user's selections to generate the next documentary chapter. The input is the user's selections and profile information, and the output is a new prompt for the generative AI model and new generated content. Another API call is made to retrieve the necessary data.
[0964] Step 9:
[0965] The generated new content is then sent to the device again, and the user begins viewing it. The input is the generated new content, and the output is data sent to the smartphone app.
[0966] Step 10:
[0967] The user views new content and answers the interactive questions again. This process is repeated between the server, the device, and the user. The input is the new interactive questions, and the output is the user's new answers.
[0968] This allows users to enjoy an endless variety of personalized documentary experiences.
[0969] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0970] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[0971] Server Operation
[0972] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[0973] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[0974] Furthermore, the emotion engine uses the user's emotional data to tailor the next content or question. This emotional data is processed in real time, so the content provided is most appropriate for the user's current emotional state.
[0975] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[0976] Device behavior
[0977] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[0978] The device is equipped with an emotion engine that analyzes the user's emotions in real time from their facial expressions, voice, etc. This emotion data is sent to the server and used to generate the next content.
[0979] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[0980] User Actions
[0981] Users log in to the platform and watch the initial documentary video provided. They make their own choices in response to interactive questions that appear while watching. The emotion engine also analyzes the user's emotions, and the emotional information is fed back to the system without any special operation.
[0982] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[0983] Specific examples
[0984] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[0985] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[0986] Additionally, the emotional state of the user while watching is analyzed by the emotion engine, and if the user is excited, for example, more detailed and visually stimulating content is provided.
[0987] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to their interests and emotions.
[0988] The system customizes documentaries according to individual users' interests and emotional state, providing an interactive viewing experience that combats viewer boredom and maximizes learning outcomes.
[0989] The processing flow will be explained below.
[0990] Server Processing
[0991] Step 1:
[0992] The server receives the user's login request and verifies the credentials.
[0993] Step 2:
[0994] The server retrieves the user's profile information from a database, which includes the user's interests, past viewing history, and individual viewing habits.
[0995] Step 3:
[0996] The server provides profile information to the generative AI model to generate the initial documentary content, which then generates the initial video based on this information.
[0997] Step 4:
[0998] The server receives the generated initial documentary video and transmits it to the terminal.
[0999] Step 5:
[1000] The server also receives data from the emotion engine to capture the user's current emotional state during video playback. The emotion data is processed in real time to influence the next content generation.
[1001] Terminal handling
[1002] Step 1:
[1003] The terminal displays an interface for the user to enter login information, which is then sent to the server.
[1004] Step 2:
[1005] The terminal plays back the initial documentary video received from the server and provides it to the user.
[1006] Step 3:
[1007] The emotion engine built into the device analyzes the user's emotions in real time from their facial expressions and voice, and this emotion data is sent to the server.
[1008] Step 4:
[1009] Interactive questions are displayed during the video, which contain multiple options for the user to select.
[1010] Step 5:
[1011] The user's selection results and emotion data are sent to the server and used as data for generating new content.
[1012] User Action
[1013] Step 1:
[1014] Users log in to the platform and watch the initial video.
[1015] Step 2:
[1016] While watching, interactive questions are displayed and the user is prompted to select one of the options.
[1017] Step 3:
[1018] The emotion engine analyzes the user's emotions in real time and sends the data to the server.
[1019] Step 4:
[1020] The user then watches the next newly generated video and answers the questions again.
[1021] Step 5:
[1022] After viewing, the user enters feedback, and this data is also sent to the server.
[1023] Specific examples
[1024] Step 1:
[1025] When a user logs in with an interest in "history," the server obtains this interest information.
[1026] Step 2:
[1027] The server requests an initial documentary related to "Ancient Egypt" from the generative AI model, while the device detects the user's emotional state, analyzing facial expressions and voice sounds that indicate excitement or curiosity in real time.
[1028] Step 3:
[1029] An initial video is generated and provided to the user. While watching the video, the user is prompted with the question, "Which part are you interested in next?" The options offered are "Pyramids," "Pharaoh's History," and "Everyday Life."
[1030] Step 4:
[1031] When the user selects "Pyramid" and the emotion engine detects the user's state of excitement, the server adjusts the next video it generates to be more visually stimulating.
[1032] Step 5:
[1033] A new video is generated and sent to the user's device, the user starts watching again, and the same process continues.
[1034] This system allows users to experience a wide variety of documentary content optimized for their emotional state.
[1035] Example 2
[1036] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1037] While conventional documentary content generation systems have the ability to generate content based on user interests and selection, they do not fully utilize user emotions or real-time feedback. This makes it difficult to fine-tune personalization based on individual user interests and emotional states, resulting in a poor viewing experience. Furthermore, it is difficult to efficiently reuse generated content, resulting in a waste of computing resources.
[1038] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1039] In this invention, the server includes a means for acquiring user profile information, a means for generating documentary content based on the user's interests and selections using a generative AI model, a means for collecting emotion data in real time using an emotion engine that recognizes the user's emotions and using that data to generate subsequent content, and a means for indexing and storing the generated content in a database. This makes it possible to provide highly personalized documentary content based on the user's interests and emotions, thereby realizing efficient use of computing resources and an improved quality of viewing experience.
[1040] "User profile information" is information that includes personal data such as a user's interests, past viewing history, preferences, and the like.
[1041] A "generative AI model" is an artificial intelligence system that generates appropriate documentary content based on a user's profile information and choices.
[1042] "Documentary content" is content consisting of video, audio and text based on facts and information.
[1043] An "interactive question" is a question with options that is presented to the user while they are watching, and the user's choice is reflected in the next content.
[1044] The "emotion engine" is a technology that collects and analyzes emotional data from users' facial expressions, voice, etc. in real time.
[1045] "Emotion data" is data that indicates the user's current emotional state, and is used to generate the next content.
[1046] A "database" is a system that efficiently stores and manages acquired information and generated content.
[1047] "Indexing" is the process of organizing data into a searchable format that makes queries within a database more efficient.
[1048] "Feedback" refers to information such as impressions and evaluations provided by users after viewing a content, and is useful for creating the next content.
[1049] "Reuse" refers to the reuse of content that has been generated once in a manner that corresponds to the selections of other users.
[1050] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[1051] Server Operation
[1052] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history. The server then provides this information to a generative AI model to generate the initial documentary content. The generative AI model can be OpenAI GPT-4 or similar.
[1053] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The generative AI model is again invoked to generate the next documentary chapter based on the user's choices, inputting their selections and profile information.
[1054] Furthermore, the system uses the user's emotional data recognized by the emotion engine to tailor the next content or question. Based on the emotional data processed in real time, the system provides content that best suits the user's current emotional state. Each generated chapter is stored and indexed in a database, allowing it to be quickly provided to other users who make the same selection, reducing computational costs.
[1055] Device behavior
[1056] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1057] The device also has a built-in emotion engine that analyzes emotions from the user's facial expressions and voice in real time and sends the data to the server. This emotion data is used to generate new content. The newly generated documentary chapter is sent to the device and provided to the user. Feedback entered by the user after watching the video is also sent from the device to the server and used to generate the next piece of content.
[1058] User Actions
[1059] Users log in to the platform and watch the initial documentary video provided to them. They then select options in response to interactive questions that appear while they are watching. Furthermore, emotional information analyzed by the emotion engine is fed back to the system in real time, so no special operations are required.
[1060] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[1061] Specific examples
[1062] For example, if a user is interested in "history," the server uses this interest information to request the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary. While watching, the question "Which part are you interested in next?" is displayed, and the options are presented: "Pyramids," "Pharaoh's history," and "Daily life." If the user selects "Pyramids," the selection is sent to the server, and the next chapter of the documentary is generated with content focused on "Pyramids."
[1063] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[1064] Prompt Sentence Examples
[1065] "If a user is interested in history, we can have the generative AI model generate a documentary on the theme of 'Ancient Egypt,' taking into account the user's specific profile information and past viewing history to provide interesting content."
[1066] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1067] Step 1:
[1068] A user logs in to the platform. The user accesses the platform and enters their username and password on the login screen. When they press the login button, the information is sent to the server. The server performs authentication, and if successful, starts the user's session.
[1069] Input: Username, Password
[1070] Output: User authentication result (success / failure), session start
[1071] Step 2:
[1072] The server retrieves the user's profile information: After successful authentication, the server retrieves the user's profile information (interests, past viewing history, etc.) from the database. This profile information is used in the next processing step.
[1073] Input: User ID
[1074] Output: User profile information
[1075] Step 3:
[1076] The server provides prompts to the generative AI model to generate initial documentary content. The server creates appropriate prompt sentences for the generative AI model based on the profile information and sends them to the model. The generative AI model generates documentary content based on these prompt sentences.
[1077] Input: User profile information, prompt text
[1078] Output: Initial documentary content
[1079] Step 4:
[1080] The server sends the generated documentary content to the device. The server receives the content generated by the generative AI model and sends it to the user's device. The content is ready for the user to watch.
[1081] Input: Early documentary content
[1082] Output: Content delivery to devices
[1083] Step 5:
[1084] The terminal displays the documentary content and the user starts viewing it. The terminal plays the content received from the server and makes it available for viewing by the user. The user starts viewing the content.
[1085] Input: Documentary content from the server
[1086] Output: Playback begins, user watches
[1087] Step 6:
[1088] The terminal displays interactive questions. While the user is watching the documentary, the terminal displays the interactive questions received from the server and presents the user with options.
[1089] Input: Interactive questions from the server
[1090] Output: Screen presenting options to the user
[1091] Step 7:
[1092] The user selects an option for an interactive question. The user selects one option from multiple options and enters it into the terminal. The terminal sends the selected information to the server.
[1093] Input: User's choice
[1094] Output: Sending choices from the terminal to the server
[1095] Step 8:
[1096] The terminal transmits the selection to the server. After receiving the user's selection, the terminal transmits the data to the server, providing input data for the next content generation.
[1097] Input: User's choice
[1098] Output: Send selected data to the server
[1099] Step 9:
[1100] The server requests a new documentary chapter from the generative AI model. The server generates a prompt for the new documentary chapter based on the user's choices and sends it to the generative AI model. The generative AI model then generates the new chapter based on that prompt.
[1101] Input: User choices, prompt text
[1102] Output: New documentary chapter
[1103] Step 10:
[1104] The device collects the user's emotional data and sends it to the server. The device's built-in emotion engine analyzes the user's emotional data while watching and sends the information to the server.
[1105] Input: User's emotional state
[1106] Output: Send emotion data to the server
[1107] Step 11:
[1108] The server generates the next content based on the emotion data. It analyzes the received emotion data and adjusts the content and questions of the next documentary chapter. The generated documentary chapter is indexed and stored in a database.
[1109] Input: User emotion data
[1110] Output: adjusted documentary chapters, stored in a database
[1111] Step 12:
[1112] The server sends new documentary chapters to the terminal, and the server sends the new chapters to the terminal so that the user can continue watching.
[1113] Enter: a new documentary chapter
[1114] Output: New documentary streaming to your device
[1115] Step 13:
[1116] User continues watching. The user watches the newly received documentary chapter and continues their viewing experience.
[1117] Enter: a new documentary chapter
[1118] Output: User continues watching
[1119] Step 14:
[1120] Users provide feedback on videos. After watching, users input their impressions and ratings into their devices, and the feedback is sent to the server. This information is used to generate new content for the next time.
[1121] Input: User feedback
[1122] Output: Send feedback to the server
[1123] The above processing steps enable the system of the present invention to provide highly personalized documentary content according to the user's interests, preferences, and emotional state.
[1124] (Application example 2)
[1125] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1126] Conventional documentary content generation systems personalize content based on user interests and selections, but they fail to take into account the user's emotional state. This makes it difficult to provide content that adapts to the user's emotional state, and it is therefore difficult to significantly improve the satisfaction of the viewing experience. Furthermore, user feedback is often not adequately reflected in the next content generation, limiting the improvement of content quality.
[1127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1128] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selection, means for indexing and storing the generated content in a database, and means for recognizing the user's emotions using an emotion engine and reflecting the emotion data in the generation of the next content, thereby enabling the provision of sophisticated personalized content according to the user's emotional state.
[1129] "User profile information" is individual information about a user, such as the user's interests and past viewing history.
[1130] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate documentary content based on user interests and choices.
[1131] An "interactive question" is a question that is presented to a user to elicit a choice from the user.
[1132] The "emotion engine" is a system that analyzes the user's facial expressions and voice to obtain emotional data.
[1133] "Documentary content" refers to fact-based video works and informational content.
[1134] "Feedback" refers to data such as reactions, evaluations, and opinions from users.
[1135] "Personalization" is a method of providing individualized experiences and content.
[1136] A "database" is a system in which information is systematically stored and can be retrieved as needed.
[1137] "Indexing" is the process of organizing information in a database so that it can be quickly searched.
[1138] The system of the present invention acquires user profile information and generates interactive documentary content using a generative AI model based on that information. By combining it with an emotion engine, it is possible to provide more sophisticated personalized content.
[1139] Server Operation
[1140] The server first retrieves the user's profile information from a database, including the user's interests and past viewing history. The server then uses this profile information to provide appropriate input data to a generative AI model to generate initial documentary content. The generated documentary content is then sent to the device, where the user can begin viewing.
[1141] Interactive questions are generated and presented to the user as they watch. To generate the next documentary chapter based on the user's choices, the AI model is again invoked, taking their selections and profile information into account. Furthermore, the emotion engine uses the user's emotional data to tailor the next content and questions. This emotional data is processed in real time, ensuring the content best suited to the user's current emotional state is presented.
[1142] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[1143] Device behavior
[1144] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1145] The device also has an emotion engine that analyzes the user's emotions in real time based on their facial expressions and voice. This emotion data is sent to the server and used to generate the next content. The newly generated documentary chapter is then sent back to the device and provided to the user.
[1146] User Actions
[1147] Users log in to the platform and watch the initial documentary video presented to them. They make their own choices in response to interactive questions displayed while watching. The emotion engine also analyzes the user's emotions, so emotional information is fed back to the system without any special operations. After watching, users provide feedback by entering their impressions and ratings into their device. This feedback is reflected in the next video generation.
[1148] Specific examples
[1149] Let's say the user is interested in "history." Based on this information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[1150] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[1151] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[1152] Hardware and software used
[1153] The server uses a cloud computing infrastructure with high-performance processors and large memory capacities. The generative AI model uses a deep learning model using the Python programming language and the TensorFlow library. The emotion engine combines facial expression analysis software and voice recognition technology. A NoSQL database is suitable for the database. The terminals are mobile devices such as smartphones, tablets, and smart glasses, which communicate with the server in real time via an internet connection.
[1154] Prompt Sentence Examples
[1155] Below are some example prompts to input to the generative AI model:
[1156] "User is interested in Ancient Egypt. Generate the next documentary chapter."
[1157] "User is interested in 'Pyramids'. Please generate the following details:"
[1158] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1159] Step 1:
[1160] The server retrieves the user's profile information from a database.
[1161] Input: User ID
[1162] Data processing and calculation: Issue database queries to extract user interests and past viewing history.
[1163] Output: User profile information
[1164] Step 2:
[1165] The server generates a prompt sentence for the generated AI model based on the profile information.
[1166] Input: User profile information
[1167] Data processing and calculation: Creates appropriate prompt statements in text format based on profile information.
[1168] Output: Prompt sentence to be passed to the generative AI model
[1169] Step 3:
[1170] The server inputs prompt sentences into the generative AI model to generate initial documentary content.
[1171] Input: prompt statement
[1172] Data processing and calculation: The generative AI model analyzes the prompt text and generates documentary content.
[1173] Output: Early documentary content
[1174] Step 4:
[1175] The server transmits the generated documentary content to the terminal.
[1176] Input: Early documentary content
[1177] Data processing and calculation: Content is optimized for the device and packetized.
[1178] Output: Data sent to the terminal
[1179] Step 5:
[1180] The terminal provides the received documentary content to the user and displays interactive questions.
[1181] Input: Early documentary content
[1182] Data processing and calculation: Play content and display interactive questions as an overlay.
[1183] Output: The content the user views and the questions
[1184] Step 6:
[1185] The user makes selections in response to interactive questions that are displayed.
[1186] Input: Interactive Question
[1187] Data processing and calculation: Obtain user selections via touch input or voice input.
[1188] Output: User's choice
[1189] Step 7:
[1190] The terminal transmits the user's selection to the server.
[1191] Input: User's choice
[1192] Data processing and calculation: The selected data is packetized for transmission to the server.
[1193] Output: Send selected data
[1194] Step 8:
[1195] The server receives the user's selection and generates the next prompt sentence along with the analysis results from the emotion engine.
[1196] Input: User choices and emotion data
[1197] Data processing and calculation: Based on the selection and emotional state, the following prompt sentence is created in text format.
[1198] Output: Next prompt statement
[1199] Step 9:
[1200] The server inputs the following prompt sentence into the generative AI model and generates the following documentary content.
[1201] Input: the following prompt statement
[1202] Data processing and calculation: The generative AI model analyzes the prompt text and generates new documentary content.
[1203] Output: Continuing documentary content
[1204] Step 10:
[1205] The server transmits the generated continuous documentary content to the terminal.
[1206] Input: Documentary content to be continued
[1207] Data processing and calculation: Content is optimized for the device and packetized.
[1208] Output: Data sent to the terminal
[1209] Step 11:
[1210] The device analyzes the user's emotional data in real time and sends the results back to the server.
[1211] Input: User's facial expression data and voice data
[1212] Data processing and calculation: Analyze using the emotion engine and obtain emotion data.
[1213] Output: Sending emotion data
[1214] Step 12:
[1215] The server stores the acquired emotional data and user feedback in a database and uses it to generate the next content.
[1216] Input: Emotion data and user feedback
[1217] Data processing and calculation: Index and store in a database.
[1218] Output: Data stored in the database
[1219] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1220] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1221] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1222] [Fourth embodiment]
[1223] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1224] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1225] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1226] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1227] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1228] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1229] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1230] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1231] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1232] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1233] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1234] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1235] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1236] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[1237] Server Operation
[1238] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[1239] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[1240] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[1241] Device behavior
[1242] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1243] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[1244] User Actions
[1245] Users log in to the platform and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that appear, which determine the content of the next video. After watching, users provide feedback that will be reflected in the next video. Through this interactive selection and feedback process, users can experience infinitely customized documentary content.
[1246] Specific examples
[1247] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[1248] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[1249] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[1250] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[1251] The processing flow will be explained below.
[1252] Step 1:
[1253] A user logs into the platform. The user enters their authentication information on the login screen and is authenticated.
[1254] Step 2:
[1255] The terminal sends the user's login information to the server, which then retrieves the user's profile information from the database based on the received authentication information.
[1256] Step 3:
[1257] The server extracts the user's interests and past viewing history from their profile information and uses that information to send an initial documentary content generation request to the generative AI model.
[1258] Step 4:
[1259] The generative AI model generates an initial documentary video and returns it to the server, which then receives the video file and sends it to the device.
[1260] Step 5:
[1261] The terminal plays the initial documentary video and provides a viewing interface to the user, who then watches the video.
[1262] Step 6:
[1263] The server generates interactive questions based on the video content and sends the questions and options to the device, which then displays them to the user.
[1264] Step 7:
[1265] The user selects one of the options for the question presented, and the selection information is sent from the terminal to the server.
[1266] Step 8:
[1267] The server receives the user's selection and sends another request to the generative AI model to generate the next documentary chapter, updating the profile information if necessary.
[1268] Step 9:
[1269] The server receives the generated video file of the next documentary chapter and sends it to the device, which plays this new video.
[1270] Step 10:
[1271] The user watches a new documentary chapter and makes selections in response to the interactive questions that are presented again, and the process is repeated.
[1272] Step 11:
[1273] The server indexes and stores each generated video in a database so that it can be quickly provided to other users who make similar selections.
[1274] Step 12:
[1275] After watching a video, the user inputs feedback into the device, which then sends the feedback to the server and uses it to generate the next content.
[1276] This allows users to experience an infinite amount of customized documentary content.
[1277] Example 1
[1278] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1279] Conventional video content is difficult to customize according to user interests, and viewers often become bored with the same content. Furthermore, while systems existed that provided interactive experiences, they lacked the ability to dynamically generate content based on user choices. As a result, the appeal of the viewing experience was limited, and learning benefits were reduced.
[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1281] In this invention, the server includes means for acquiring user profile information, means for generating interactive content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selections, means for indexing and storing each generated chapter in a database, and means for reusing the generated content to quickly provide it to other users who make the same selections, thereby enabling dynamic and infinite generation of customized documentary content according to the user's interests.
[1282] "User profile information" is data that indicates the individual attributes and behavioral patterns of each user, such as the user's interests, past viewing history, and selection history.
[1283] "Generative AI model" refers to an artificial intelligence algorithm or framework for dynamically generating documentary content based on user profile information and selection information.
[1284] An "interactive question" is a question that is displayed while the user is watching, prompting the user to make a choice, and is an element that influences the generation of the next content.
[1285] "Indexing" is the process of organizing and classifying generated content into a database so that it can be efficiently searched and referenced.
[1286] "Reuse" refers to reducing the computational cost of the system by making content once generated available to other users who make the same or similar choices.
[1287] "Feedback" refers to ratings and comments provided by users after viewing a content, and is sent to the server as a reference for the next content generation.
[1288] "Terminal" refers to an apparatus or device through which a user views content and inputs selections and feedback.
[1289] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.
[1290] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history, which serves as the basic data for content generation. The server then generates documentary content using a generative AI model (for example, OpenAI's GPT-4 or TensorFlow). The prompt uses a format such as "Generate the content of the first interactive documentary based on the user's profile information." The generated content is then sent to the device.
[1291] The user uses a device to watch this initial documentary content. While watching, the server generates interactive questions, sends them to the device, and presents them to the user. For example, in response to the question, "Which part are you interested in next?", options such as "Pyramids," "Pharaoh's History," and "Daily Life" are presented. The user's selection is sent from the device to the server, and the next content is generated based on this. Again, the prompt is in the form, "Based on the user's selection and profile information, generate the next chapter of the documentary about the pyramids."
[1292] Each generated chapter is stored and indexed in a database on the server, allowing it to be quickly provided to other users who make the same selection, reducing computational costs. The content generated through reuse also corresponds to other users' selections.
[1293] After watching, users can enter their feedback into their device, which then sends it to the server. The server then incorporates this feedback into future content generation. For example, if a user gives a high rating to the "Pyramid" chapter, the server will take this into account in future content generation and adjust the content generation to provide more detailed information.
[1294] As a concrete example, if a user is interested in "history," the server generates a documentary on the theme of "ancient Egypt" based on the user's profile information. The prompt text used is, "Based on the user's profile information, generate an interactive documentary about ancient Egypt." The generated video is displayed on the device, and the user begins watching. Interactive questions are presented during viewing, and the next documentary chapter is generated based on the user's selection. In this way, an infinite documentary experience completely customized to the user's interests is realized.
[1295] This system allows users to enjoy an interactive and dynamic viewing experience by providing documentary content customized to their interests, thereby preventing user boredom and maximizing learning effectiveness.
[1296] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1297] Step 1:
[1298] The server receives the user's identification information (user ID) when the user logs into the platform. As input, the user ID is sent to the server. The server uses this user ID to retrieve the corresponding profile information (interests, viewing history, etc.) from the database. As output, the retrieved profile information is ready to be passed to the generative AI model.
[1299] Step 2:
[1300] The server uses the acquired profile information to create a prompt appropriate for the generative AI model. This prompt includes instructions for generating initial documentary content based on the user's interests. For example, a prompt might be created such as, "Generate an interactive documentary about ancient Egypt based on the user's profile information." The acquired profile information is used as input, and the prompt is generated as output.
[1301] Step 3:
[1302] The server inputs the generated prompt sentences into a generative AI model to generate initial documentary content. The generative AI model generates data such as text, images, and videos based on the input prompt sentences. The output is the initial documentary content, which includes content that is in line with the user's interests.
[1303] Step 4:
[1304] The server sends the generated initial documentary content to the terminal, which uses the generated documentary content as input and transfers it to the terminal as output, in an appropriate format (e.g., video file or streaming data).
[1305] Step 5:
[1306] The terminal displays the initial documentary content received from the server to the user, who then begins viewing the content, using the documentary content received from the server as input and providing viewing to the user as output.
[1307] Step 6:
[1308] The server generates interactive questions according to the viewing progress. For example, when a specific scene in a video is reached, the server generates a question such as "Which part are you interested in next?" and presents options such as "Pyramids," "Pharaoh's history," and "Daily life." The viewing progress data is used as input, and the interactive questions are generated as output.
[1309] Step 7:
[1310] The terminal presents the user with interactive questions sent from the server. The user selects one of the options presented. The user's option is entered into the terminal as input, and the selection is sent to the server as output.
[1311] Step 8:
[1312] The server sends a prompt to the generative AI model again based on the user's selection and profile information to generate the next chapter of documentary content. The prompt might be something like, "Generate the next chapter of a documentary about pyramids based on the user's selection and profile information." The user's selection and profile information are used as input, and the documentary content for the next chapter is generated as output.
[1313] Step 9:
[1314] The server sends the generated documentary content of the next chapter to the terminal. The terminal receives this new content and provides it to the user. The generated documentary content is used as input and sent to the terminal as output.
[1315] Step 10:
[1316] A user watches new documentary content and then provides feedback. As input, the feedback information is input to the terminal, and as output, the feedback is sent to the server.
[1317] Step 11:
[1318] The server stores the feedback received from the device in a database and uses it for the next content generation. The user's feedback information is used as input and stored in the database as output. This feedback information is used to adjust the prompt text to be generated next time.
[1319] Step 12:
[1320] The server analyzes the accumulated feedback and reflects it in the next content generation. This allows for more highly customized content based on the user's interests and satisfaction. The feedback data is used as input, and the output is used to adjust the generated prompts and the generative AI model.
[1321] (Application example 1)
[1322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1323] Conventional content delivery systems have difficulty generating and providing personalized content based on user interests and preferences, resulting in inconsistent viewing experiences. Furthermore, the generated content lacks interactivity, making it difficult to maintain user interest over the long term. Furthermore, there is a lack of a way to utilize user feedback to improve the quality of the viewing experience.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1325] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, and means for delivering the generated content via a smartphone application and collecting user feedback, thereby enabling the generation of personalized content based on the user's interests, providing an interactive experience, and utilizing the collected feedback for future content generation.
[1326] "User profile information" is information including the user's interests, past viewing history, personal data, etc., and is the basis for content generation.
[1327] A "generative AI model" is a model that uses artificial intelligence to generate content based on user profile information and interactive choices.
[1328] "Documentary content" refers to content that is comprised of fact-based video and audio, and includes personalized content tailored to the user's interests.
[1329] An "interactive question" is a choice or question presented to generate different content based on a user's selection.
[1330] A "smartphone application" is an application that runs on a smartphone and provides an interface for users to answer interactive questions and view content.
[1331] "Feedback" refers to the act of a user providing opinions and impressions about content they have viewed, and the information provided can be used to help generate the next piece of content.
[1332] "Indexing" is the process of organizing generated content and storing it in a database for easy searching and reuse.
[1333] A system for implementing this invention acquires user profile information and uses that information to generate interactive documentary content based on the user's interests and preferences using a generative AI model. The system is composed of the following components:
[1334] 1. Server Operation
[1335] The server retrieves the user's profile information from a database, including their interests and past viewing history. Based on this profile information, the server provides appropriate input data to a generative AI model to generate the initial documentary content. The generative AI model used is OpenAI GPT-3, a general artificial intelligence platform. The generated content is then sent to the device, where the user can begin watching.
[1336] 2. Device Operation
[1337] The user's smartphone application receives the initial documentary content from the server. When the user begins watching, the application displays interactive questions. The user's selections in response to these questions are sent from the device to the server. To generate the next documentary chapter, the server again invokes the generative AI model, inputting the user's selections and profile information. The newly generated documentary chapter is again sent to the device and served to the user.
[1338] 3. User Actions
[1339] Users log in to the smartphone application and watch the initial documentary video presented to them. While watching, they make their own choices in response to interactive questions that determine the content of the next video. After watching, users provide feedback, and that information is reflected in the next content generation.
[1340] Specific examples
[1341] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[1342] While watching, the question "Which part are you interested in next?" is displayed, and the options are "Pyramids," "Pharaoh's History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focusing on "Pyramids." Next, the following sentence is used as an example of a prompt sentence for the generative AI model:
[1343] Example prompt sentence:
[1344] "Username is interested in history, particularly ancient Egypt. Please generate a 10-minute documentary on the subject."
[1345] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to the user's interests.
[1346] The system customizes documentaries according to individual user interests and provides an interactive viewing experience, thereby combating viewer boredom and maximizing learning outcomes.
[1347] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1348] Step 1:
[1349] The server retrieves user profile information from a database. The input is a user ID or other identifying information, and the output is profile information including the user's interests and past viewing history. It filters and retrieves the required data from the database (e.g., MySQL) using SQL queries.
[1350] Step 2:
[1351] The server generates a content generation prompt for the generative AI model based on the acquired profile information. The input is the user profile information, and the output is a prompt sentence for the generative AI model. Specifically, the server analyzes the profile information using a Python script and constructs an appropriate prompt sentence.
[1352] Step 3:
[1353] The server inputs a prompt into a generative AI model (e.g., OpenAI GPT-3) to generate the initial documentary content. The input is the prompt, and the output is the generated documentary content. Specifically, the server makes an API call and receives the generated text and video data.
[1354] Step 4:
[1355] The server sends the generated content to the user's smartphone app. The input is the generated documentary content, and the output is data transmission to the smartphone app. Data is sent using an HTTP request.
[1356] Step 5:
[1357] The device (smartphone app) displays the received documentary content to the user. The user begins watching and is presented with interactive questions. The input is the content received from the server, and the output is the interface displayed to the user. Specifically, the display uses UI elements created with React Native or Flutter.
[1358] Step 6:
[1359] The user answers the interactive questions presented to them. The input is the interactive question, and the output is the user's answer. When the user presses the selection button, the selection is registered within the application.
[1360] Step 7:
[1361] The terminal sends the user's answer information to the server. The input is the user's answer information, and the output is data sent to the server. The data is sent using an HTTP POST request.
[1362] Step 8:
[1363] The server uses the user's selections to generate the next documentary chapter. The input is the user's selections and profile information, and the output is a new prompt for the generative AI model and new generated content. Another API call is made to retrieve the necessary data.
[1364] Step 9:
[1365] The generated new content is then sent to the device again, and the user begins viewing it. The input is the generated new content, and the output is data sent to the smartphone app.
[1366] Step 10:
[1367] The user views new content and answers the interactive questions again. This process is repeated between the server, the device, and the user. The input is the new interactive questions, and the output is the user's new answers.
[1368] This allows users to enjoy an endless variety of personalized documentary experiences.
[1369] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1370] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[1371] Server Operation
[1372] The server first retrieves the user's profile information from a database, including their interests and past viewing history. The server then uses this profile information to provide appropriate input data to the generative AI model to generate the initial documentary content.
[1373] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The AI model is again invoked to generate the next documentary chapter based on the user's choices and profile information.
[1374] Furthermore, the emotion engine uses the user's emotional data to tailor the next content or question. This emotional data is processed in real time, so the content provided is most appropriate for the user's current emotional state.
[1375] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[1376] Device behavior
[1377] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1378] The device is equipped with an emotion engine that analyzes the user's emotions in real time from their facial expressions, voice, etc. This emotion data is sent to the server and used to generate the next content.
[1379] The newly generated documentary chapter is then sent back to the device and provided to the user. After watching the video, the user can enter their feedback on the video into the device. The device then sends this feedback to the server and uses it to generate the next content.
[1380] User Actions
[1381] Users log in to the platform and watch the initial documentary video provided. They make their own choices in response to interactive questions that appear while watching. The emotion engine also analyzes the user's emotions, and the emotional information is fed back to the system without any special operation.
[1382] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[1383] Specific examples
[1384] Let's say the user is interested in "history." Based on this interest information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[1385] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[1386] Additionally, the emotional state of the user while watching is analyzed by the emotion engine, and if the user is excited, for example, more detailed and visually stimulating content is provided.
[1387] New content is generated and sent to the user's device, and the user starts watching again, and the process continues, providing a limitless documentary experience tailored to their interests and emotions.
[1388] The system customizes documentaries according to individual users' interests and emotional state, providing an interactive viewing experience that combats viewer boredom and maximizes learning outcomes.
[1389] The processing flow will be explained below.
[1390] Server Processing
[1391] Step 1:
[1392] The server receives the user's login request and verifies the credentials.
[1393] Step 2:
[1394] The server retrieves the user's profile information from a database, which includes the user's interests, past viewing history, and individual viewing habits.
[1395] Step 3:
[1396] The server provides profile information to the generative AI model to generate the initial documentary content, which then generates the initial video based on this information.
[1397] Step 4:
[1398] The server receives the generated initial documentary video and transmits it to the terminal.
[1399] Step 5:
[1400] The server also receives data from the emotion engine to capture the user's current emotional state during video playback. The emotion data is processed in real time to influence the next content generation.
[1401] Terminal handling
[1402] Step 1:
[1403] The terminal displays an interface for the user to enter login information, which is then sent to the server.
[1404] Step 2:
[1405] The terminal plays back the initial documentary video received from the server and provides it to the user.
[1406] Step 3:
[1407] The emotion engine built into the device analyzes the user's emotions in real time from their facial expressions and voice, and this emotion data is sent to the server.
[1408] Step 4:
[1409] Interactive questions are displayed during the video, which contain multiple options for the user to select.
[1410] Step 5:
[1411] The user's selection results and emotion data are sent to the server and used as data for generating new content.
[1412] User Action
[1413] Step 1:
[1414] Users log in to the platform and watch the initial video.
[1415] Step 2:
[1416] While watching, interactive questions are displayed and the user is prompted to select one of the options.
[1417] Step 3:
[1418] The emotion engine analyzes the user's emotions in real time and sends the data to the server.
[1419] Step 4:
[1420] The user then watches the next newly generated video and answers the questions again.
[1421] Step 5:
[1422] After viewing, the user enters feedback, and this data is also sent to the server.
[1423] Specific examples
[1424] Step 1:
[1425] When a user logs in with an interest in "history," the server obtains this interest information.
[1426] Step 2:
[1427] The server requests an initial documentary related to "Ancient Egypt" from the generative AI model, while the device detects the user's emotional state, analyzing facial expressions and voice sounds that indicate excitement or curiosity in real time.
[1428] Step 3:
[1429] An initial video is generated and provided to the user. While watching the video, the user is prompted with the question, "Which part are you interested in next?" The options offered are "Pyramids," "Pharaoh's History," and "Everyday Life."
[1430] Step 4:
[1431] When the user selects "Pyramid" and the emotion engine detects the user's state of excitement, the server adjusts the next video it generates to be more visually stimulating.
[1432] Step 5:
[1433] A new video is generated and sent to the user's device, the user starts watching again, and the same process continues.
[1434] This system allows users to experience a wide variety of documentary content optimized for their emotional state.
[1435] Example 2
[1436] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1437] While conventional documentary content generation systems have the ability to generate content based on user interests and selection, they do not fully utilize user emotions or real-time feedback. This makes it difficult to fine-tune personalization based on individual user interests and emotional states, resulting in a poor viewing experience. Furthermore, it is difficult to efficiently reuse generated content, resulting in a waste of computing resources.
[1438] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1439] In this invention, the server includes a means for acquiring user profile information, a means for generating documentary content based on the user's interests and selections using a generative AI model, a means for collecting emotion data in real time using an emotion engine that recognizes the user's emotions and using that data to generate subsequent content, and a means for indexing and storing the generated content in a database. This makes it possible to provide highly personalized documentary content based on the user's interests and emotions, thereby realizing efficient use of computing resources and an improved quality of viewing experience.
[1440] "User profile information" is information that includes personal data such as a user's interests, past viewing history, preferences, and the like.
[1441] A "generative AI model" is an artificial intelligence system that generates appropriate documentary content based on a user's profile information and choices.
[1442] "Documentary content" is content consisting of video, audio and text based on facts and information.
[1443] An "interactive question" is a question with options that is presented to the user while they are watching, and the user's choice is reflected in the next content.
[1444] The "emotion engine" is a technology that collects and analyzes emotional data from users' facial expressions, voice, etc. in real time.
[1445] "Emotion data" is data that indicates the user's current emotional state, and is used to generate the next content.
[1446] A "database" is a system that efficiently stores and manages acquired information and generated content.
[1447] "Indexing" is the process of organizing data into a searchable format that makes queries within a database more efficient.
[1448] "Feedback" refers to information such as impressions and evaluations provided by users after viewing a content, and is useful for creating the next content.
[1449] "Reuse" refers to the reuse of content that has been generated once in a manner that corresponds to the selections of other users.
[1450] The system of the present invention consists of a server that acquires user profile information and generates interactive documentary content using a generative AI model based on that information, a terminal that plays the generated content and receives selections and feedback from the user, and a user that provides selections and feedback.Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more sophisticated personalized content.
[1451] Server Operation
[1452] The server first retrieves the user's profile information from a database. This profile information includes the user's interests and past viewing history. The server then provides this information to a generative AI model to generate the initial documentary content. The generative AI model can be OpenAI GPT-4 or similar.
[1453] The generated documentary content is sent to the device, where the user begins watching. Interactive questions are generated and presented to the user during viewing. The generative AI model is again invoked to generate the next documentary chapter based on the user's choices, inputting their selections and profile information.
[1454] Furthermore, the system uses the user's emotional data recognized by the emotion engine to tailor the next content or question. Based on the emotional data processed in real time, the system provides content that best suits the user's current emotional state. Each generated chapter is stored and indexed in a database, allowing it to be quickly provided to other users who make the same selection, reducing computational costs.
[1455] Device behavior
[1456] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1457] The device also has a built-in emotion engine that analyzes emotions from the user's facial expressions and voice in real time and sends the data to the server. This emotion data is used to generate new content. The newly generated documentary chapter is sent to the device and provided to the user. Feedback entered by the user after watching the video is also sent from the device to the server and used to generate the next piece of content.
[1458] User Actions
[1459] Users log in to the platform and watch the initial documentary video provided to them. They then select options in response to interactive questions that appear while they are watching. Furthermore, emotional information analyzed by the emotion engine is fed back to the system in real time, so no special operations are required.
[1460] After watching, users can provide feedback by entering their impressions and ratings into the device, which will be reflected in the next video generation.
[1461] Specific examples
[1462] For example, if a user is interested in "history," the server uses this interest information to request the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary. While watching, the question "Which part are you interested in next?" is displayed, and the options are presented: "Pyramids," "Pharaoh's history," and "Daily life." If the user selects "Pyramids," the selection is sent to the server, and the next chapter of the documentary is generated with content focused on "Pyramids."
[1463] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[1464] Prompt Sentence Examples
[1465] "If a user is interested in history, we can have the generative AI model generate a documentary on the theme of 'Ancient Egypt,' taking into account the user's specific profile information and past viewing history to provide interesting content."
[1466] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1467] Step 1:
[1468] A user logs in to the platform. The user accesses the platform and enters their username and password on the login screen. When they press the login button, the information is sent to the server. The server performs authentication, and if successful, starts the user's session.
[1469] Input: Username, Password
[1470] Output: User authentication result (success / failure), session start
[1471] Step 2:
[1472] The server retrieves the user's profile information: After successful authentication, the server retrieves the user's profile information (interests, past viewing history, etc.) from the database. This profile information is used in the next processing step.
[1473] Input: User ID
[1474] Output: User profile information
[1475] Step 3:
[1476] The server provides prompts to the generative AI model to generate initial documentary content. The server creates appropriate prompt sentences for the generative AI model based on the profile information and sends them to the model. The generative AI model generates documentary content based on these prompt sentences.
[1477] Input: User profile information, prompt text
[1478] Output: Initial documentary content
[1479] Step 4:
[1480] The server sends the generated documentary content to the device. The server receives the content generated by the generative AI model and sends it to the user's device. The content is ready for the user to watch.
[1481] Input: Early documentary content
[1482] Output: Content delivery to devices
[1483] Step 5:
[1484] The terminal displays the documentary content and the user starts viewing it. The terminal plays the content received from the server and makes it available for viewing by the user. The user starts viewing the content.
[1485] Input: Documentary content from the server
[1486] Output: Playback begins, user watches
[1487] Step 6:
[1488] The terminal displays interactive questions. While the user is watching the documentary, the terminal displays the interactive questions received from the server and presents the user with options.
[1489] Input: Interactive questions from the server
[1490] Output: Screen presenting options to the user
[1491] Step 7:
[1492] The user selects an option for an interactive question. The user selects one option from multiple options and enters it into the terminal. The terminal sends the selected information to the server.
[1493] Input: User's choice
[1494] Output: Sending choices from the terminal to the server
[1495] Step 8:
[1496] The terminal transmits the selection to the server. After receiving the user's selection, the terminal transmits the data to the server, providing input data for the next content generation.
[1497] Input: User's choice
[1498] Output: Send selected data to the server
[1499] Step 9:
[1500] The server requests a new documentary chapter from the generative AI model. The server generates a prompt for the new documentary chapter based on the user's choices and sends it to the generative AI model. The generative AI model then generates the new chapter based on that prompt.
[1501] Input: User choices, prompt text
[1502] Output: New documentary chapter
[1503] Step 10:
[1504] The device collects the user's emotional data and sends it to the server. The device's built-in emotion engine analyzes the user's emotional data while watching and sends the information to the server.
[1505] Input: User's emotional state
[1506] Output: Send emotion data to the server
[1507] Step 11:
[1508] The server generates the next content based on the emotion data. It analyzes the received emotion data and adjusts the content and questions of the next documentary chapter. The generated documentary chapter is indexed and stored in a database.
[1509] Input: User emotion data
[1510] Output: adjusted documentary chapters, stored in a database
[1511] Step 12:
[1512] The server sends new documentary chapters to the terminal, and the server sends the new chapters to the terminal so that the user can continue watching.
[1513] Enter: a new documentary chapter
[1514] Output: New documentary streaming to your device
[1515] Step 13:
[1516] User continues watching. The user watches the newly received documentary chapter and continues their viewing experience.
[1517] Enter: a new documentary chapter
[1518] Output: User continues watching
[1519] Step 14:
[1520] Users provide feedback on videos. After watching, users input their impressions and ratings into their devices, and the feedback is sent to the server. This information is used to generate new content for the next time.
[1521] Input: User feedback
[1522] Output: Send feedback to the server
[1523] The above processing steps enable the system of the present invention to provide highly personalized documentary content according to the user's interests, preferences, and emotional state.
[1524] (Application example 2)
[1525] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1526] Conventional documentary content generation systems personalize content based on user interests and selections, but they fail to take into account the user's emotional state. This makes it difficult to provide content that adapts to the user's emotional state, and it is therefore difficult to significantly improve the satisfaction of the viewing experience. Furthermore, user feedback is often not adequately reflected in the next content generation, limiting the improvement of content quality.
[1527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1528] In this invention, the server includes means for acquiring user profile information, means for generating documentary content based on the user's interests and selections using a generative AI model, means for providing the generated content to the user and generating and presenting interactive questions, means for regenerating the next documentary content based on the user's selection, means for indexing and storing the generated content in a database, and means for recognizing the user's emotions using an emotion engine and reflecting the emotion data in the generation of the next content, thereby enabling the provision of sophisticated personalized content according to the user's emotional state.
[1529] "User profile information" is individual information about a user, such as the user's interests and past viewing history.
[1530] A "generative AI model" is an algorithm or system that uses artificial intelligence to generate documentary content based on user interests and choices.
[1531] An "interactive question" is a question that is presented to a user to elicit a choice from the user.
[1532] The "emotion engine" is a system that analyzes the user's facial expressions and voice to obtain emotional data.
[1533] "Documentary content" refers to fact-based video works and informational content.
[1534] "Feedback" refers to data such as reactions, evaluations, and opinions from users.
[1535] "Personalization" is a method of providing individualized experiences and content.
[1536] A "database" is a system in which information is systematically stored and can be retrieved as needed.
[1537] "Indexing" is the process of organizing information in a database so that it can be quickly searched.
[1538] The system of the present invention acquires user profile information and generates interactive documentary content using a generative AI model based on that information. By combining it with an emotion engine, it is possible to provide more sophisticated personalized content.
[1539] Server Operation
[1540] The server first retrieves the user's profile information from a database, including the user's interests and past viewing history. The server then uses this profile information to provide appropriate input data to a generative AI model to generate initial documentary content. The generated documentary content is then sent to the device, where the user can begin viewing.
[1541] Interactive questions are generated and presented to the user as they watch. To generate the next documentary chapter based on the user's choices, the AI model is again invoked, taking their selections and profile information into account. Furthermore, the emotion engine uses the user's emotional data to tailor the next content and questions. This emotional data is processed in real time, ensuring the content best suited to the user's current emotional state is presented.
[1542] Each generated chapter is stored and indexed in a database so that it can be quickly provided to other users who make the same selection, reducing computational costs through reuse. This process is repeated as long as the user continues to make selections.
[1543] Device behavior
[1544] The device receives the initial documentary content from the server when the user logs in to the platform. When the user starts watching, interactive questions are displayed. The user's selections are sent from the device to the server and wait until the next video is generated.
[1545] The device also has an emotion engine that analyzes the user's emotions in real time based on their facial expressions and voice. This emotion data is sent to the server and used to generate the next content. The newly generated documentary chapter is then sent back to the device and provided to the user.
[1546] User Actions
[1547] Users log in to the platform and watch the initial documentary video presented to them. They make their own choices in response to interactive questions displayed while watching. The emotion engine also analyzes the user's emotions, so emotional information is fed back to the system without any special operations. After watching, users provide feedback by entering their impressions and ratings into their device. This feedback is reflected in the next video generation.
[1548] Specific examples
[1549] Let's say the user is interested in "history." Based on this information, the server requests the generative AI model to generate a documentary on the theme of "ancient Egypt." The generated video is displayed on the device, and the user watches the documentary.
[1550] While watching, the user is asked, "Which part are you interested in next?" and is presented with the options "Pyramids," "Pharaonic History," and "Daily Life." If the user selects "Pyramids," the selection is sent to the server, and the next documentary chapter is generated with content focused on "Pyramids."
[1551] Furthermore, the emotion engine analyzes the user's emotional state while watching. For example, if the user is excited, more detailed and visually stimulating content is provided. New content is generated and sent to the user's device. The user then resumes watching, and the same process continues. This allows users to enjoy an infinitely expanding documentary experience tailored to their interests and emotions.
[1552] Hardware and software used
[1553] The server uses a cloud computing infrastructure with high-performance processors and large memory capacities. The generative AI model uses a deep learning model using the Python programming language and the TensorFlow library. The emotion engine combines facial expression analysis software and voice recognition technology. A NoSQL database is suitable for the database. The terminals are mobile devices such as smartphones, tablets, and smart glasses, which communicate with the server in real time via an internet connection.
[1554] Prompt Sentence Examples
[1555] Below are some example prompts to input to the generative AI model:
[1556] "User is interested in Ancient Egypt. Generate the next documentary chapter."
[1557] "User is interested in 'Pyramids'. Please generate the following details:"
[1558] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1559] Step 1:
[1560] The server retrieves the user's profile information from a database.
[1561] Input: User ID
[1562] Data processing and calculation: Issue database queries to extract user interests and past viewing history.
[1563] Output: User profile information
[1564] Step 2:
[1565] The server generates a prompt sentence for the generated AI model based on the profile information.
[1566] Input: User profile information
[1567] Data processing and calculation: Creates appropriate prompt statements in text format based on profile information.
[1568] Output: Prompt sentence to be passed to the generative AI model
[1569] Step 3:
[1570] The server inputs prompt sentences into the generative AI model to generate initial documentary content.
[1571] Input: prompt statement
[1572] Data processing and calculation: The generative AI model analyzes the prompt text and generates documentary content.
[1573] Output: Early documentary content
[1574] Step 4:
[1575] The server transmits the generated documentary content to the terminal.
[1576] Input: Early documentary content
[1577] Data processing and calculation: Content is optimized for the device and packetized.
[1578] Output: Data sent to the terminal
[1579] Step 5:
[1580] The terminal provides the received documentary content to the user and displays interactive questions.
[1581] Input: Early documentary content
[1582] Data processing and calculation: Play content and display interactive questions as an overlay.
[1583] Output: The content the user views and the questions
[1584] Step 6:
[1585] The user makes selections in response to interactive questions that are displayed.
[1586] Input: Interactive Question
[1587] Data processing and calculation: Obtain user selections via touch input or voice input.
[1588] Output: User's choice
[1589] Step 7:
[1590] The terminal transmits the user's selection to the server.
[1591] Input: User's choice
[1592] Data processing and calculation: The selected data is packetized for transmission to the server.
[1593] Output: Send selected data
[1594] Step 8:
[1595] The server receives the user's selection and generates the next prompt sentence along with the analysis results from the emotion engine.
[1596] Input: User choices and emotion data
[1597] Data processing and calculation: Based on the selection and emotional state, the following prompt sentence is created in text format.
[1598] Output: Next prompt statement
[1599] Step 9:
[1600] The server inputs the following prompt sentence into the generative AI model and generates the following documentary content.
[1601] Input: the following prompt statement
[1602] Data processing and calculation: The generative AI model analyzes the prompt text and generates new documentary content.
[1603] Output: Continuing documentary content
[1604] Step 10:
[1605] The server transmits the generated continuous documentary content to the terminal.
[1606] Input: Documentary content to be continued
[1607] Data processing and calculation: Content is optimized for the device and packetized.
[1608] Output: Data sent to the terminal
[1609] Step 11:
[1610] The device analyzes the user's emotional data in real time and sends the results back to the server.
[1611] Input: User's facial expression data and voice data
[1612] Data processing and calculation: Analyze using the emotion engine and obtain emotion data.
[1613] Output: Sending emotion data
[1614] Step 12:
[1615] The server stores the acquired emotional data and user feedback in a database and uses it to generate the next content.
[1616] Input: Emotion data and user feedback
[1617] Data processing and calculation: Index and store in a database.
[1618] Output: Data stored in the database
[1619] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1620] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1621] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1622] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1623] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1624] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1625] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1626] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1627] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1628] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1629] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1630] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1631] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1632] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1633] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1634] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1635] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1636] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1637] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1638] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1639] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1640] The following is further disclosed regarding the above embodiment.
[1641] (Claim 1)
[1642] means for obtaining user profile information;
[1643] a means for generating documentary content based on user interests and selections using a generative AI model;
[1644] means for providing the generated content to a user and generating and presenting interactive questions;
[1645] means for regenerating the next documentary content based on the user's selection;
[1646] a means for indexing and storing the generated content in a database;
[1647] A system including:
[1648] (Claim 2)
[1649] 10. The system according to claim 1, further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
[1650] (Claim 3)
[1651] 10. The system of claim 1, further comprising means for reusing and saving the generated documentary content in a reusable form to accommodate other user selections.
[1652] "Example 1"
[1653] (Claim 1)
[1654] means for obtaining user profile information;
[1655] A means for generating interactive content based on user interests and choices using a generative AI model; and
[1656] means for providing the generated content to a user and generating and presenting interactive questions;
[1657] means for regenerating the next documentary content based on the user's selection;
[1658] A means of indexing and storing each generated chapter in a database;
[1659] a means for reusing the generated content to quickly serve other users who make the same selection;
[1660] A system including:
[1661] (Claim 2)
[1662] 10. The system according to claim 1, further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
[1663] (Claim 3)
[1664] 10. The system of claim 1, further comprising means for storing the generated interactive content in a reusable form and reusing it to accommodate other user selections.
[1665] "Application Example 1"
[1666] (Claim 1)
[1667] means for obtaining user profile information;
[1668] a means for generating documentary content based on user interests and selections using a generative AI model;
[1669] means for providing the generated content to a user and generating and presenting interactive questions;
[1670] means for regenerating the next documentary content based on the user's selection;
[1671] means for delivering the generated content via a smartphone application and collecting feedback from users;
[1672] a means for indexing and storing the generated content in a database;
[1673] A system including:
[1674] (Claim 2)
[1675] 10. The system according to claim 1, further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
[1676] (Claim 3)
[1677] 10. The system of claim 1, further comprising means for reusing and saving the generated documentary content in a reusable form to accommodate other user selections.
[1678] "Example 2: Combining Emotion Engines"
[1679] (Claim 1)
[1680] means for obtaining user profile information;
[1681] a means for generating documentary content based on user interests and selections using a generative AI model;
[1682] means for providing the generated content to a user and generating and presenting interactive questions;
[1683] means for regenerating the next documentary content based on the user's selection;
[1684] A means for collecting emotional data in real time using an emotion engine that recognizes the user's emotions and using the data to generate the next content;
[1685] a means for indexing and storing the generated content in a database;
[1686] A system including:
[1687] (Claim 2)
[1688] 10. The system according to claim 1, further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
[1689] (Claim 3)
[1690] 10. The system of claim 1, further comprising means for reusing and saving the generated documentary content in a reusable form to accommodate other user selections.
[1691] "Application example 2 when combining emotion engines"
[1692] (Claim 1)
[1693] means for obtaining user profile information;
[1694] a means for generating documentary content based on user interests and selections using a generative AI model;
[1695] means for providing the generated content to a user and generating and presenting interactive questions;
[1696] means for regenerating the next documentary content based on the user's selection;
[1697] a means for indexing and storing the generated content in a database;
[1698] A means for recognizing a user's emotions using an emotion engine and reflecting the emotion data in the next content generation;
[1699] A system including:
[1700] (Claim 2)
[1701] 10. The system according to claim 1, further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
[1702] (Claim 3)
[1703] 10. The system of claim 1, further comprising means for reusing and saving the generated documentary content in a reusable form to accommodate other user selections. [Explanation of symbols]
[1704] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for obtaining user profile information; a means for generating documentary content based on user interests and selections using a generative AI model; means for providing the generated content to a user and generating and presenting interactive questions; means for regenerating the next documentary content based on the user's selection; a means for indexing and storing the generated content in a database; A system including:
2. The system according to claim 1 , further comprising means for collecting feedback from the user's terminal and utilizing the feedback in next content generation.
3. 10. The system of claim 1, further comprising means for reusing and saving the generated documentary content in a reusable form to accommodate other user selections.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A