system
The system addresses the limitations of modern social media by allowing users to generate and share personalized video content of life events with privacy controls, facilitating easy sharing and diverse styles.
Patent Information
- Application Number
- JP2024138677
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Modern social media platforms limit the means by which users can visualize and share specific events or turning points in their lives, making it difficult to share important events with others and lacking sufficient privacy management and flexible visibility settings.
A system that allows users to input information about events, generate video content using AI, set disclosure scopes, and share it with others while respecting privacy, enabling users to create and share videos in various styles.
Enables users to easily visualize and share important events in diverse styles while protecting privacy, allowing flexible disclosure settings.
Smart Images

Figure 2026036162000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Modern social media platforms limit the means by which users can visualize and share specific events or turning points in their lives. As a result, users cannot easily share important events in their lives with others, and it is difficult to receive others' reactions to those events. Furthermore, privacy management for shared information is often insufficient, and the visibility settings are often unclear. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system including: a means for a user to input information about events and turning points that occurred in the user's life; a generating device for generating video content based on the input information; a display device for displaying the generated video content; a means for setting a disclosure scope; and a means for making the video content available to other users according to the disclosure scope. This system allows users to create videos of specific events in their lives and easily share them with others while respecting privacy. It is also possible to upload multiple materials (photos, audio, music, etc.) and generate video content in a variety of styles based on them.
[0006] "User" refers to an individual who uses the system to provide information about events and turning points in their life.
[0007] "Generator" refers to a combination of hardware and software for generating video content based on user-entered information.
[0008] "Display Device" means a device for visually presenting generated video content to a user or other users.
[0009] "Means for setting the scope of disclosure" refers to an interface and related functions that allow a user to make settings to limit the viewers of the generated video content.
[0010] "Means for disclosing video content to other users according to the disclosure range" refers to a function for displaying video content to relevant users based on the set disclosure range.
[0011] "Materials" refers to digital data such as photos, audio, and music provided by users.
[0012] "Analysis means" refers to the function for analyzing uploaded materials and extracting the information necessary for video generation.
[0013] "Image" refers to the style and atmosphere of the video content (e.g., anime-style, drama-style, American comic book-style, etc.).
[0014] "Style change means" refers to a function for changing the appearance and atmosphere of video content based on an image selected by the user.
[0015] "Video Content" means visual and audio content generated from information and materials provided by users. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention relates to a system in which users input information about events and turning points that occurred in their lives, and AI generates video content based on that information and shares it with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[0038] System configuration
[0039] 1. User login
[0040] The user is presented with an interface to log into the system, where they enter their login information and are authenticated.
[0041] 2. Conducting AI interviews
[0042] After logging in, the server checks the user's profile information and begins the AI interview.
[0043] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[0044] 3. Provision of Materials
[0045] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[0046] The server stores and analyzes the uploaded material.
[0047] 4. Select output format
[0048] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[0049] Users can also select additional options such as trends and historical background of the time.
[0050] 5. Image generation and confirmation
[0051] The server uses AI algorithms to generate video content based on user input and provided materials.
[0052] Users can check the generated video on the preview page and request corrections if necessary.
[0053] 6. Setting the visibility and publishing
[0054] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[0055] The server makes the video content available to other users and manages access rights based on the settings.
[0056] Specific examples
[0057] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0058] 1. The user logs in and the AI interview begins.
[0059] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[0060] 3. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[0061] 4. The server analyzes the user's material and recommends a video style (for example, anime style) based on that analysis.
[0062] 5. The user selects an anime style and then selects the "Trend of the Year" option.
[0063] 6. The server generates the video content based on the selected style and options.
[0064] 7. The user reviews the generated footage and requests corrections if necessary.
[0065] 8. The server completes the edited video.
[0066] 9. The user selects "Private" and sets the video content to be shared only with family members.
[0067] 10. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[0068] The system allows users to record important events in their lives and easily share them with others while respecting privacy.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] A user visits the system's website or app and enters their login information.
[0072] Step 2:
[0073] The server authenticates the login information and verifies the user's profile information.
[0074] Step 3:
[0075] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[0076] Step 4:
[0077] The user enters text responses to the AI interview questions.
[0078] Step 5:
[0079] The server analyzes the user's answers and generates and displays the next question (e.g., "What is the duration of the event?").
[0080] Step 6:
[0081] The user answers additional questions as the interview continues.
[0082] Step 7:
[0083] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[0084] Step 8:
[0085] The server receives the uploaded material and stores it in storage.
[0086] Step 9:
[0087] The server analyzes the stored material and extracts the information necessary to generate the video.
[0088] Step 10:
[0089] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[0090] Step 11:
[0091] The user selects their preferred style and options (such as what was popular at the time).
[0092] Step 12:
[0093] The server stores the user's selections and prepares the video for generation.
[0094] Step 13:
[0095] The server uses AI algorithms to generate video content based on user input and provided materials.
[0096] Step 14:
[0097] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[0098] Step 15:
[0099] The user can check the footage on the preview page and enter correction requests if necessary.
[0100] Step 16:
[0101] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[0102] Step 17:
[0103] The user makes a final confirmation and sets the public range (private, limited, public).
[0104] Step 18:
[0105] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] In modern society, there are limited ways to record special events and important moments in video and share them with others. As a result, users face challenges in easily creating and sharing video content, requiring significant effort and technical knowledge. Furthermore, there is a lack of flexible ways to change the style and image of video content, making it difficult to meet the diverse needs of users. Furthermore, from the perspective of privacy protection, a lack of systems that allow users to easily set and manage the scope of disclosure has been pointed out.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes means for a user to input information about events or important events in his or her life, means for generating and presenting questions to the user and collecting the information input by the user, means for uploading multiple digital data provided by the user, means for analyzing and saving the uploaded digital data, means for a generation device to generate video content based on the collected information and analyzed data, means for a display device to display the generated video content, means for setting a disclosure range, and means for making the video content available to other users in accordance with the disclosure range. This enables users, without special knowledge, to easily visualize their important events, express them in various styles, and share them with others while protecting their privacy.
[0111] A "user" is someone who uses the system to visualize their own events or important events and share them with others.
[0112] "Server" refers to the central computer that controls and manages the entire system, and analyzes and stores information entered by users and uploaded material data.
[0113] "Terminal" means a device that allows a user to access the system, input information, and view video content.
[0114] A "generation device" is a device or software that uses AI algorithms to generate video content based on information collected from users and materials provided by them.
[0115] A "display device" is a device for displaying generated video content to a user.
[0116] The "means for generating questions" is a function for generating appropriate questions based on information input by the user and presenting them to the user.
[0117] "Digital data" refers to multimedia materials such as photographs, audio, and music provided by users.
[0118] "Means for uploading" refers to the functionality that allows users to transfer digital data to the system.
[0119] "Means for analyzing and storing" refers to the function for analyzing uploaded digital data and storing it within the system.
[0120] The "means for setting the disclosure range" is a function for selecting and setting the disclosure range of the generated video content.
[0121] The "means for publishing according to the public range" is a function for making video content accessible to other users based on the public range that has been set.
[0122] This invention relates to a system in which users input information about events or important events in their lives, and AI generates video content based on that information and shares it with other users. This system mainly involves information exchange and processing between a server, a terminal, and the user.
[0123] System configuration
[0124] 1. User login
[0125] To access the system, the terminal displays a login screen. The user enters their email address and password, which the terminal sends to the server. The server authenticates them by referencing a database (e.g., MongoDB or MySQL®), and if authentication is successful, the user is allowed to proceed to the next step.
[0126] 2. Conducting AI interviews
[0127] The server retrieves the profile information of the user who has successfully logged in and generates appropriate questions using a natural language processing model (e.g., GPT-4 (registered trademark)). The server sends a series of questions to the device, which displays them to the user. The user answers the questions, and the device sends the answers to the server.
[0128] 3. Provision of Materials
[0129] Users use their devices to upload digital data such as their photos, audio files, and songs. The devices send this data to a server, which then analyzes the data using image processing libraries (e.g., OpenCV) and audio analysis libraries (e.g., Librosa) and stores it in a storage system (e.g., Amazon S3).
[0130] 4. Select output format
[0131] The device displays an interface for selecting the style and image of the video content, allowing the user to choose styles such as anime, TV drama, American comics, etc. Additional options such as the trends of the time and historical background can also be selected. The device sends the selection results to the server, which stores them in a database.
[0132] 5. Image generation and confirmation
[0133] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials. The server temporarily saves the generated video and sends a URL to the device. The device displays a video preview screen, which the user can check. If the user sends a correction request as needed, the server accepts the correction request and regenerates the video.
[0134] 6. Setting the visibility and publishing
[0135] The device displays an interface for setting the visibility of the video content. The user can select from options such as private, limited, or public. The device sends the selection result to the server, which then manages access rights for the video content based on the selected visibility.
[0136] Specific examples
[0137] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0138] 1. The user logs in and the system starts the AI interview.
[0139] 2. The server sends the user a question such as, "Which event from your graduation ceremony was most memorable?", and the user answers with a specific episode.
[0140] 3. The user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and their favorite songs from their device.
[0141] 4. The server analyzes the uploaded material and suggests a video style (e.g., anime style).
[0142] 5. The user selects an anime style and chooses "Trend of the Year" as an additional option.
[0143] 6. The server generates the video content based on the selected style and options, using DeepArt or StyleGAN.
[0144] 7. The user checks the generated video on their device and requests corrections if necessary.
[0145] 8. The server completes the edited video.
[0146] 9. The user selects "Private" and sets the video content to be shared only with family members.
[0147] 10. The server publishes the video based on the user's settings and manages it so that only users within the specified range of access can access the video.
[0148] Example prompts to input to the generative AI model
[0149] A possible prompt for a generative AI model (e.g., GPT-4) might look something like this:
[0150] The user is talking about their college graduation ceremony. Use the following elements to generate animated video content:
[0151] Photo 1: A photo of me with my friends at the graduation ceremony.
[0152] Photo 2: Exterior of the ceremony venue.
[0153] Audio file: Audio of the speech on graduation day.
[0154] Song: A memorable song from graduation ceremony.
[0155] The tone of the video should be emotive and reflect the trends of the time.
[0156] This system allows users, without any special knowledge, to easily visualize their important events, express them in a variety of styles, and share them with others while protecting their privacy.
[0157] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0158] System program processing flow
[0159] Step 1:
[0160] User login
[0161] The device displays the login screen.
[0162] Input: User's email address and password.
[0163] Specific operation: Display a login form using HTML / CSS / JavaScript (registered trademark).
[0164] The user enters their email address and password.
[0165] Output: The login information entered.
[0166] The terminal sends the input information to the server.
[0167] Input: The login information entered.
[0168] Output: HTTP request to the server.
[0169] The server performs authentication by referencing a database (e.g. MongoDB or MySQL).
[0170] Input: The login information sent to the server.
[0171] Specific behavior: Provides an API (e.g., RESTful API) on the backend and handles authentication logic.
[0172] Output: Authentication result (success / failure).
[0173] The server returns the authentication result to the terminal, and if successful, proceeds to the next step.
[0174] Input: Authentication result.
[0175] Output: Next screen based on authentication result (instructions to proceed to next step if successful).
[0176] Step 2:
[0177] Conducting AI interviews
[0178] The server retrieves the user's profile information after a successful login.
[0179] Input: Login success notification.
[0180] Specific Actions: Retrieve user profile from database.
[0181] Output: Profile information.
[0182] The server uses the profile information to generate appropriate questions using a natural language processing model (e.g., GPT-4).
[0183] Input: Profile information.
[0184] What it does: Generate questions using GPT-4.
[0185] Output: Question list.
[0186] The server sends the generated question to the terminal.
[0187] Input: Questionnaire.
[0188] Output: HTTP response containing the questionnaire.
[0189] The device displays a question and the user enters the answer.
[0190] Input: Questionnaire.
[0191] Specific operation: Dynamically generate and display a question form.
[0192] Output: The user's answer.
[0193] The device sends the user's answer to the server.
[0194] Input: The user's answer.
[0195] Output: HTTP request to the server (response data).
[0196] Step 3:
[0197] Provision of materials
[0198] The device will display the interface for uploading materials.
[0199] Input: The upload request.
[0200] Specific behavior: HTML <input type="file"> Implement file selection using an element.
[0201] Output: Upload screen.
[0202] Users select and upload digital data such as photos, audio, and music.
[0203] Input: Materials (photos, audio, music, etc.).
[0204] Output: The digital data to be uploaded.
[0205] The device sends the uploaded data to the server.
[0206] Input: Digital data to be uploaded.
[0207] Output: HTTP request to server (digital data).
[0208] The server receives the data and stores it in a repository.
[0209] Input: Uploaded digital data.
[0210] Specific behavior: Stores data in a file storage system (e.g., Amazon S3).
[0211] Output: Stored digital data.
[0212] The server analyzes the stored data.
[0213] Input: Stored digital data.
[0214] Specific operation: Analyze the data using an image processing library (e.g., OpenCV) or a sound analysis library (e.g., Librosa).
[0215] Output: Analysis results.
[0216] Step 4:
[0217] Selecting the output format
[0218] The terminal displays an interface for selecting the style of the video content.
[0219] Input: A style selection request.
[0220] What it does: Provide users with choices using drop-down menus or radio buttons.
[0221] Output: Style selection screen.
[0222] The user selects a style and image from the options provided.
[0223] Enter: Style selection.
[0224] Output: Selected style information.
[0225] The terminal transmits the selection result to the server.
[0226] Input: Selected style information.
[0227] Output: HTTP request to the server (style information).
[0228] The server stores the user's selection in a database.
[0229] Input: Selected style information.
[0230] Specific action: Save to database.
[0231] Output: The saved style information.
[0232] Step 5:
[0233] Image generation and confirmation
[0234] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials.
[0235] Input: User input information and materials provided.
[0236] Specific operation: Generate images using DeepArt and StyleGAN.
[0237] Output: The generated video content.
[0238] The server temporarily stores the generated video and sends its URL to the terminal.
[0239] Input: Generated video content.
[0240] Output: URL of the archived video.
[0241] The device displays a video preview screen, and the user checks the video.
[0242] Input: The URL of the downloaded video.
[0243] Specific behavior: HTML <video>Display video using tags.
[0244] Output: The confirmation status of the user.
[0245] The user sends a modification request to the server as needed.
[0246] Input: The user's correction request.
[0247] Output: HTTP request to the server (modification request).
[0248] Step 6:
[0249] Setting the visibility and publishing
[0250] The device displays an interface for selecting the disclosure range.
[0251] Input: Disclosure scope setting request.
[0252] Specific behavior: Provide options for the scope of disclosure using radio buttons or checkboxes.
[0253] Output: Public range selection screen.
[0254] The user chooses the scope of disclosure.
[0255] Input: Select the disclosure range.
[0256] Output: Selected disclosure information.
[0257] The terminal transmits the selection result to the server.
[0258] Input: Selected disclosure information.
[0259] Output: HTTP request to the server (public range information).
[0260] The server sets the video content to be made public based on the selected public range.
[0261] Input: Selected disclosure information.
[0262] Specific operation: Set the scope of disclosure using ACL (Access Control List).
[0263] Output: The set disclosure range.
[0264] The server manages access rights for other users based on the public settings.
[0265] Input: The set disclosure range.
[0266] Output: Managed access rights.
[0267] (Application example 1)
[0268] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0269] In today's virtual stores, it is difficult to provide a personalized shopping experience based on users' emotions and personal memories. Therefore, a system that allows users to enjoy shopping while feeling an emotional connection is needed. Furthermore, existing systems lack the means to effectively utilize user-provided materials, display them as video content, and provide product suggestions and special offers. Furthermore, the technology to visualize individual user events and turning points and provide them as moving experiences within the virtual store is underdeveloped.
[0270] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0271] In this invention, the server includes: means for a user to input information about events and turning points that occurred in the user's life; means for a generation device to generate video content based on the input information; means for a display device to display the generated video content; means for setting a disclosure range; means for disclosing the video content to other users according to the disclosure range; and means for displaying products and special offers using the generated video content so that the user can personalize an emotional shopping experience in the virtual store. This allows the user to enjoy a virtual shopping experience based on personal memories while feeling an emotional connection, and further allows the user to effectively receive product suggestions and special offers through the video content.
[0272] "User" refers to an individual who uses the System.
[0273] "Events" refer to memorable experiences or turning points that occurred in the user's life.
[0274] A "turning point" refers to a time or event in a user's life when an important change or decision was made.
[0275] "Means for inputting information" refers to the interface through which a user provides information about their event to the system.
[0276] "Generation device" refers to a computer program and hardware system that creates video content based on input information.
[0277] "Display device" refers to a device that visually presents generated video content to a user.
[0278] The "means for setting the public range" refers to an interface that allows a user to select the viewing range of the generated video content.
[0279] "Publicity" refers to the range of users who can view the generated video content.
[0280] "Virtual store" refers to a virtual shopping environment provided on the Internet.
[0281] "Shopping Experience" refers to the series of activities in which a user browses, selects, and purchases products within a virtual store.
[0282] "Means for displaying products and special offers" refers to an interface for presenting related products and discount information within the generated video content.
[0283] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content and displays products and special offers in a virtual store. Specific embodiments of the system are described below.
[0284] System configuration
[0285] 1. User login
[0286] Users use the interface to log in to the system. When the user enters their login information, the server performs authentication.
[0287] 2. Conducting AI interviews
[0288] After logging in, the server checks the user's profile information and begins an AI interview. The server sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[0289] 3. Provision of Materials
[0290] Users upload digital data such as photos, audio, and favorite music through their devices, and the server stores and analyzes the uploaded material.
[0291] 4. Select output format
[0292] The server provides an interface for users to select the style and image of the video content (e.g., anime, TV drama, American comic book, etc.) Users can also select additional options such as the trends of the time and the historical background.
[0293] 5. Video Generation and Application to Virtual Stores
[0294] The server uses a generative AI model to generate video content based on user input and provided materials, which is then used to display products and special offers within the virtual store.
[0295] Hardware and software used
[0296] Hardware: smartphones, tablets, servers
[0297] Software: Online platform APIs, generative AI models (e.g., OpenAI® GPT-3®)
[0298] Data processing: Analyzes user input data (text, images, audio) and performs processes ranging from prompt generation to video content generation.
[0299] Specific examples
[0300] For example, consider the case where a user wants to create a video of their "college graduation ceremony." After logging in, the user answers questions such as "Which event from your graduation ceremony was most memorable?" through an AI interview. The server analyzes the materials provided by the user, such as photos, audio files, and favorite songs, and recommends video styles based on these. If the user selects an anime-style style and also selects "Trend of the Year" as an option, the server uses a generative AI model to generate video content based on the selected style and options. This video is then used to display products and special offers within the virtual store.
[0301] Prompt Sentence Examples
[0302] text
[0303] Title: College Graduation
[0304] Details: Wonderful days spent with many friends, and a graduation day that will remain in my memory forever.
[0305] Image files: ["graduation_photo1.jpg", "graduation_photo2.jpg"]
[0306] Audio file: ["graduation_speech.mp3"]
[0307] Style: Emotional documentary
[0308] Generate effective footage.
[0309] In this way, a system is realized that provides an emotional shopping experience in a virtual store based on the user's special events.
[0310] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0311] Step 1:
[0312] A user logs in to the system.
[0313] Input: User login information (username, password).
[0314] Output: If authentication is successful, proceed to next step. If not, prompt to retry.
[0315] Specific operation: The user enters the username and password into the login interface, and the server authenticates it. If the authentication is successful, the server obtains the user's profile information and proceeds to the next step.
[0316] Step 2:
[0317] AI interview begins.
[0318] Input: User profile information and login state.
[0319] Output: Detailed information about the user's events and turning points.
[0320] What it does: After logging in, the server asks the user a series of questions, which the user answers, and the server collects these answers, recording details of events and turning points that the user enters during this process.
[0321] Step 3:
[0322] Users upload materials.
[0323] Input: Digital data provided by the user, such as photos, audio, or favorite songs.
[0324] Output: Uploaded material is stored on the server and analyzed.
[0325] How it works: Users use their devices to upload photos, audio files, music, and other content to a server. The server then analyzes and stores the data. Image recognition and audio analysis algorithms are used for the analysis.
[0326] Step 4:
[0327] Video content style selection.
[0328] Input: User's event details and uploaded materials.
[0329] Output: User-selected video style and additional options.
[0330] Specific operation: Through the interface, the server allows the user to select the style of the video content (e.g., anime, drama, American comics, etc.). Additional options include the selection of the trends of the time and the historical background.
[0331] Step 5:
[0332] Video content generation.
[0333] Input: User event details, uploaded footage, selected video style and additional options.
[0334] Output: The generated video content.
[0335] Specific operation: The server uses the generative AI model to generate video content based on the user's input information and materials, inputting prompt sentences into the generative AI model and outputting video data based on them.
[0336] Step 6:
[0337] Displaying products and special offers in a virtual store.
[0338] Input: Generated video content.
[0339] Output: Products and special offers displayed in a virtual store with emotional video content.
[0340] What it does: The generated video content is played in the virtual store, where product suggestions and special offers related to the user's current situation are displayed, allowing the user to browse and purchase products with an emotional connection.
[0341] Step 7:
[0342] Setting the visibility of video content and publishing it.
[0343] Input: Generated video content, user-selected visibility.
[0344] Output: Video content with access rights controlled for other users based on the set visibility.
[0345] Specific operation: The server provides an interface for users to set the visibility of video content (private, limited, public, etc.). Based on the visibility set by the user, the server makes the video available to other users and manages access rights.
[0346] Through these steps, the system can visualize the user's special events and provide an emotional shopping experience within the virtual store.
[0347] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0348] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[0349] System configuration
[0350] 1. User login
[0351] A user visits the system's website or app and enters their login information.
[0352] 2. Conducting AI interviews
[0353] After logging in, the server checks the user's profile information and begins the AI interview.
[0354] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[0355] 3. Emotion recognition using the emotion engine
[0356] The server activates an emotion engine to recognize the user's emotion from the content of the user's answers and the provided materials.
[0357] The emotion engine analyzes the user's emotional state based on the context, words used, tone of voice, etc.
[0358] 4. Provision of Materials
[0359] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[0360] The server stores and analyzes the uploaded material.
[0361] 5. Select output format
[0362] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[0363] The server proposes an appropriate video style based on the results of the emotion engine.
[0364] 6. Image generation and confirmation
[0365] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[0366] Users can check the generated video on the preview page and request corrections if necessary.
[0367] 7. Setting the visibility and publishing
[0368] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[0369] The server makes the video content available to other users and manages access rights based on the settings.
[0370] 8. User response monitoring
[0371] The server monitors the user's reaction to the generated video content and collects emotional information about the user.
[0372] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[0373] Specific examples
[0374] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0375] 1. The user logs in and the AI interview begins.
[0376] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[0377] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[0378] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[0379] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[0380] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[0381] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[0382] 8. The user reviews the generated footage and requests corrections if necessary.
[0383] 9. The server completes the edited video.
[0384] 10. The user selects "Private" and sets the video content to be shared only with family members.
[0385] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[0386] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[0387] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[0388] The processing flow will be explained below.
[0389] Step 1:
[0390] A user visits the system's website or app and enters their login information.
[0391] Step 2:
[0392] The server authenticates the login information and verifies the user's profile information.
[0393] Step 3:
[0394] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[0395] Step 4:
[0396] The user enters text responses to the AI interview questions.
[0397] Step 5:
[0398] The server receives the user's response and activates the emotion engine to analyze the user's emotional state. During this process, emotions are recognized from the content and context of the user's response, the words used, and the tone of voice.
[0399] Step 6:
[0400] The server stores the analysis results of the emotion engine and generates and displays the next question (e.g., "What is the duration of the event?").
[0401] Step 7:
[0402] The user answers additional questions as the interview continues.
[0403] Step 8:
[0404] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[0405] Step 9:
[0406] The server receives the uploaded material and stores it in storage.
[0407] Step 10:
[0408] The server analyzes the stored material and extracts the information necessary to generate the video.
[0409] Step 11:
[0410] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[0411] Step 12:
[0412] The user selects their preferred style and options (such as what was popular at the time).
[0413] Step 13:
[0414] The server presents a suggested output format (e.g., an emotional anime style) based on the analysis results of the emotion engine.
[0415] Step 14:
[0416] The server stores the user's selections and prepares the video for generation.
[0417] Step 15:
[0418] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[0419] Step 16:
[0420] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[0421] Step 17:
[0422] The user can check the footage on the preview page and enter correction requests if necessary.
[0423] Step 18:
[0424] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[0425] Step 19:
[0426] The user makes a final confirmation and sets the public range (private, limited, public).
[0427] Step 20:
[0428] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[0429] Step 21:
[0430] The server monitors viewers' reactions to the published video content and collects user emotional information.
[0431] Step 22:
[0432] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[0433] Example 2
[0434] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0435] Today's users have a growing need to digitize the events and turning points in their lives and save and share them as emotionally charged video content. However, existing systems lacked the technology to properly analyze users' emotions and reflect them in the video content. Furthermore, they lacked the means to generate different types of video content and allow users to easily set the visibility of their content, which hindered the user experience.
[0436] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input information about events and turning points that occurred in the user's life, a means for the server to generate video content using an AI algorithm based on the input information, and a means for an emotion engine to recognize emotions from the user's responses and provided materials. This enables a user to generate video content that reflects their own emotions, select a video style with a different image, easily set the scope of publication, and share it with other users.
[0437] A "user" is an individual or organization that uses the system and provides information about events or turning points that have occurred in their life.
[0438] "Input means" refers to the interface and devices that allow users to input information about events and turning points in their lives into the system.
[0439] "Generation means" refers to the function and device by which the server generates video content using an AI algorithm based on input information.
[0440] "Emotion Engine" refers to software and systems for recognizing and analyzing emotions from user responses and provided materials.
[0441] The "display means" refers to a device and interface for visually presenting the generated video content to the user.
[0442] The "publication range setting means" is an interface and device for allowing a user to set the publicity range of video content created by the user.
[0443] The "publication means" refers to a function and device that makes video content public to other users according to a set publicity range.
[0444] "Materials" is a general term for digital data such as photographs, audio data, and music provided by users.
[0445] "Uploading means" refers to the interface and device through which users send materials from their own terminals to the system.
[0446] "Analysis means" refers to a function and device that analyzes uploaded materials and stores them in a database.
[0447] The "image selection means" refers to an interface and device that allows the user to select different video styles (e.g., anime style, drama style, American comic style, etc.) when generating video content.
[0448] "Style change means" refers to functions and devices that automatically change the style of video content based on user selection.
[0449] The present invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, a server uses an AI algorithm to generate video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. The system is implemented as follows.
[0450] The system works through the website or application that the user accesses. The user first enters their email address and password on the login screen and is authenticated by the server. Once authenticated, the user is redirected to an interface where they can take an AI interview.
[0451] The server checks the user's profile information and uses a generative AI model (e.g., a general-purpose conversation model) to send the user a series of questions. When the user answers questions such as "What was the most memorable event during your time at university?", the server stores the answers in a database and passes them to an emotion engine. The emotion engine uses Google's (registered trademark) emotion analysis API or similar to analyze the context and tone of the answers to identify the user's emotional state.
[0452] Next, users upload materials (digital data such as photos, audio files, and music) from their devices through the interface. The devices then send the materials to the server, which analyzes and stores them. Image recognition and audio analysis technologies are used for the analysis.
[0453] Next, the server suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred style from the presented styles. The selected style and options are saved on the server and used for video generation.
[0454] The server combines the selected style, emotional state, and user-provided materials to generate video content using a generative AI model (e.g., DeepMotion's AI animation generation technology). The user can view the generated video on a preview page and request corrections if necessary. The server applies the corrections and regenerates the video.
[0455] Finally, the server displays an interface for setting the visibility of the video content, and the user selects the visibility. The server then makes the video available to other users based on the settings and manages access rights. In this case, the server manages access rights for the video content according to the visibility, allowing for limited or general visibility.
[0456] In addition, the server monitors viewers' reactions to the video content after it has been released, reanalyzing the collected emotional information and saving it as reference information for the next video. This process allows users to create video content that reflects their own emotions and share it with other users within appropriate public limits.
[0457] As a concrete example, consider the case where a user wants to create a video of their college graduation ceremony. First, the user logs in, and the server sends a question such as, "Which event from your graduation ceremony was most memorable?" After the user answers with a specific episode, the server performs sentiment analysis to identify the user's emotions (e.g., joy, excitement, nostalgia). Next, the user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and favorite songs. The server analyzes these materials and recommends video styles based on the sentiment analysis results. The user selects a suggested style, and the server uses AI technology to generate video content based on the selected style and options. The user reviews the generated video and makes any necessary adjustments. Finally, the user selects "private" and shares the video content only with family members. The server manages the video based on the settings, monitors viewer reactions, and uses them to generate the next video. In this way, users can easily create and share emotionally relevant videos of their important events.
[0458] An example of a prompt sentence is asking the user, "Which event from your college graduation ceremony was the most memorable?" In this way, detailed information about the user's experience is collected, and emotion analysis and video generation are performed based on that information.
[0459] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0460] Step 1:
[0461] A user accesses the system's website or app and enters their email address and password on the login screen. The server compares this information with the database, and if authentication is successful, the user is taken to the dashboard screen.
[0462] Input: User's email address and password
[0463] Data processing: Check the input information against the database
[0464] Output: User authentication result, dashboard screen displayed if authentication is successful
[0465] Step 2:
[0466] After logging in, the server checks the user's profile information and uses the generative AI model to initiate an AI interview, sending the user a series of questions to which they respond.
[0467] Input: User profile information, AI model prompt
[0468] Data processing: Generate appropriate questions based on AI algorithms
[0469] Output: Sending interview questions to the user
[0470] Step 3:
[0471] The user answers questions sent by the server. The answers are entered as text or voice. The server receives these answers and stores them in a database.
[0472] Input: User response (text or voice)
[0473] Data processing: Saving response data
[0474] Output: Response data stored in a database
[0475] Step 4:
[0476] The server passes the user's response data to the emotion engine, which uses natural language processing technology to analyze the context and tone of the response and identify the emotional state (e.g., joy, sadness, surprise).
[0477] Input: User response data
[0478] Data processing: Sentiment analysis (using NLP technology)
[0479] Output: Identified emotional state data
[0480] Step 5:
[0481] Users upload digital materials such as photos, audio files, and music from their own devices through the interface. The devices send this data to the server, which then analyzes and stores it in a database.
[0482] Input: User photos, audio files, songs
[0483] Data processing: Analysis and storage of material data
[0484] Output: Material data stored in a database
[0485] Step 6:
[0486] The server displays an interface that suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred video style from the presented options.
[0487] Input: User sentiment analysis results, existing video style data
[0488] Data processing: Proposing appropriate video styles
[0489] Output: User suggestions, user choices
[0490] Step 7:
[0491] The server integrates the selected style, emotional state, and user-provided materials to generate video content using a generative AI model, and the user can view the generated video content on a preview page.
[0492] Input: Selected visual style, emotional state, user material data
[0493] Data processing: Image generation using AI algorithms
[0494] Output: Generated video content
[0495] Step 8:
[0496] The user checks the generated video content and requests corrections if necessary. The server receives the correction request and regenerates the video.
[0497] Input: User's correction request
[0498] Data processing: Regenerating a modified version of the video
[0499] Output: Modified video content
[0500] Step 9:
[0501] The server displays an interface for setting the visibility of video content, and the user selects a setting from options such as private, limited, or public. Based on the setting, the video is made public to other users and access rights are managed.
[0502] Input: User's visibility settings
[0503] Data processing: Applying disclosure settings
[0504] Output: Published video content
[0505] Step 10:
[0506] After the video content is released, the server monitors the viewer's reaction to it, reanalyzes it using the emotion engine, and saves the information as reference for the next video generation.
[0507] Input: Viewer response data (likes, comments, etc.)
[0508] Data processing: sentiment analysis, data storage
[0509] Output: Reference data for future video generation
[0510] In this way, users can create and share video content that reflects their own emotions.
[0511] (Application example 2)
[0512] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0513] In today's world, people have a growing desire to record and share important events and turning points in their lives as video. However, this process requires a lot of time and effort, and it is particularly challenging to generate video that expresses appropriate emotions for each event. Furthermore, determining the extent to which the generated video should be shared and monitoring viewer reactions to collect emotions are also complex. A simple and effective method to solve these problems is needed.
[0514] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input information about events and turning points that occurred in the user's life, means for a generation device to generate video content based on the input information, means for a display device to display the generated video content, means for setting a disclosure range, means for disclosing the video content to other users in accordance with the disclosure range, and means for monitoring viewer reactions and collecting user emotional information. This enables a user to easily generate video content that reflects the emotions of important events, share it within an appropriate disclosure range, and further collect viewer reactions to use in generating the next video.
[0515] "User" refers to an individual or group that uses the system to input information about events and turning points that have occurred in their life, and to generate and share video content.
[0516] "Generation device" refers to a device or software that automatically generates video content based on information input by a user.
[0517] A "display device" is a device or interface for visually displaying generated video content to a user.
[0518] "Means for setting the scope of disclosure" refers to a function or interface that allows users to select and set to whom and to what extent the generated video content will be disclosed.
[0519] "Viewers" refer to other users (third parties) who view published video content.
[0520] "Monitoring" refers to the process of observing and recording audience reactions and behavior, especially capturing audience emotional responses.
[0521] "Emotional information" refers to data or analytical results that indicate the emotional state of a user or viewer.
[0522] "Multiple Materials" refers to digital media files such as photos, audio, and music uploaded by users.
[0523] A "generative AI model" is an artificial intelligence technology that generates a video script based on information input by the user and emotional analysis.
[0524] A "prompt sentence" is an input text given to a generative AI model to instruct it on the content and style of the video content to be generated.
[0525] This invention relates to a system in which users input information about events and turning points in their lives, and AI generates video content based on that information, and an emotion engine is used to recognize and share the user's emotions. This system involves a series of processes, including user information input, preview of the generated video, setting the sharing range, and monitoring viewer reactions.
[0526] System configuration
[0527] User login
[0528] The server provides a means for users to access the system's website or application and enter their login information, which receives the user's authentication information and logs them into the system.
[0529] Conducting AI interviews
[0530] After logging in, the server checks the user's profile information and begins an AI interview. Specifically, it sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[0531] Emotion recognition with emotion engine
[0532] The server then activates an emotion engine based on the user's responses to recognize the user's emotions. The emotion engine analyzes the user's emotional state from their words and tone of voice.
[0533] Provision of materials and analysis
[0534] Users upload materials for video content, such as photos, audio, and music, from their devices. The server stores and analyzes the uploaded materials.
[0535] Selecting the output format
[0536] Users can use an interface to select the style of the video content (e.g., anime, drama, comic, etc.). The server then suggests an appropriate video style based on the results of the emotion engine.
[0537] Image generation and confirmation
[0538] The server uses a generative AI model to generate video content based on user input, provided materials, and recognized emotions. Users can view the generated video on a preview page and request corrections if necessary.
[0539] Setting the visibility and publishing
[0540] The server displays an interface for setting the visibility of the video content, and the user can select the visibility (e.g., private, limited, public). The server makes the video content available to other users based on the settings and manages access rights.
[0541] User response monitoring
[0542] The server monitors the viewer's reaction to the generated video content and collects information on the user's emotions, which is then saved as reference for the next video generation.
[0543] Hardware and software used
[0544] The system is implemented using the following hardware and software:
[0545] Servers: Web servers, database servers
[0546] Device: PC, smartphone, tablet, etc. that users access
[0547] Software: Generative AI models (e.g., OpenAI), emotion analysis engines (e.g., emotion_recognition library)
[0548] Specific examples
[0549] For example, a user may want to film their "college graduation ceremony."
[0550] procedure
[0551] 1. The user logs in and the AI interview begins.
[0552] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[0553] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[0554] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[0555] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[0556] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[0557] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[0558] 8. The user reviews the generated footage and requests corrections if necessary.
[0559] 9. The server completes the edited video.
[0560] 10. The user selects "Private" and sets the video content to be shared only with family members.
[0561] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[0562] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[0563] Prompt Sentence Examples
[0564] For example, the following prompt sentence is input to the generative AI model:
[0565] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[0566] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[0567] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0568] Step 1:
[0569] A user visits the system's website or application and enters their login information. The input to this step is the user's credentials (e.g., email address and password), and the output is a message indicating authentication success or failure. The server checks the user's credentials against information in its database and, if there is a match, allows the user to log in.
[0570] Step 2:
[0571] After logging in, the server checks the user's profile information and starts the AI interview. The input is the user's profile information and a pre-registered question list, and the output is sending questions to the user. The server sequentially sends the user questions such as "Which event at your graduation ceremony was most memorable?"
[0572] Step 3:
[0573] The user answers questions from the server. The input of this step is the user's text answer, and the output is the collected episode details. The user answers specific episodes based on their own experiences, and the information is sent to the server.
[0574] Step 4:
[0575] The server activates an emotion engine based on the collected responses to recognize the user's emotions. The input of this step is the user's response information, and the output is analyzed emotion data. The emotion engine analyzes the language tone and keywords from the user's response to identify the user's emotional state.
[0576] Step 5:
[0577] Users upload materials for video content, such as photos, audio, and music, from their devices to the server. The input is the digital media files (photos, audio, and music) uploaded by the user, and the output is a message confirming the completion of the upload. The device sends the provided materials to the server, which then stores them.
[0578] Step 6:
[0579] The server analyzes the uploaded material and proposes a style for the video content based on it. The input of this step is the uploaded material and emotional data, and the output is a recommended video style. The server analyzes the user's material and emotional state and recommends an appropriate style (e.g., an emotional anime style).
[0580] Step 7:
[0581] The user selects a suggested video style and then selects further options. The input is the style suggestion from the server and the user's selection, and the output is the selected style and options. The user makes their selection through the interface, and that information is sent to the server.
[0582] Step 8:
[0583] The server generates video content using a generative AI model based on the selected style and options. The input for this step is the user's selection information and emotional data, which are converted into specific prompt sentences. The output is the generated video content. The server sends the following prompt sentence to the generative AI model:
[0584] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[0585] Step 9:
[0586] The user can check the generated video on the preview page and request corrections if necessary. The input is the generated video content, and the output is the user's feedback or correction requests. The server displays the generated video to the user and receives any corrections that need to be made.
[0587] Step 10:
[0588] The server completes the revised video. The input of this step is the user's feedback, and the output is the completed video content. The server regenerates the video based on the user's feedback and provides the final version.
[0589] Step 11:
[0590] The user sets the visibility of the completed video content. The input is the visibility selection, and the output is a confirmation message that the settings have been completed. The server centrally makes the video available to other users based on the visibility (private, limited, public) entered.
[0591] Step 12:
[0592] The server monitors viewers' reactions to video content and collects emotional information. The input is viewer reaction data, and the output is analyzed emotional information. The server monitors viewers' reactions and comments and records them as reference information for the next video generation.
[0593] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0594] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0595] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0596] [Second embodiment]
[0597] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0598] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0599] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0600] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0601] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0602] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0603] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0604] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0605] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0606] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0607] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0608] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0609] This invention relates to a system in which users input information about events and turning points that occurred in their lives, and AI generates video content based on that information and shares it with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[0610] System configuration
[0611] 1. User login
[0612] The user is presented with an interface to log into the system, where they enter their login information and are authenticated.
[0613] 2. Conducting AI interviews
[0614] After logging in, the server checks the user's profile information and begins the AI interview.
[0615] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[0616] 3. Provision of Materials
[0617] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[0618] The server stores and analyzes the uploaded material.
[0619] 4. Select output format
[0620] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[0621] Users can also select additional options such as trends and historical background of the time.
[0622] 5. Image generation and confirmation
[0623] The server uses AI algorithms to generate video content based on user input and provided materials.
[0624] Users can check the generated video on the preview page and request corrections if necessary.
[0625] 6. Setting the visibility and publishing
[0626] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[0627] The server makes the video content available to other users and manages access rights based on the settings.
[0628] Specific examples
[0629] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0630] 1. The user logs in and the AI interview begins.
[0631] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[0632] 3. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[0633] 4. The server analyzes the user's material and recommends a video style (for example, anime style) based on that analysis.
[0634] 5. The user selects an anime style and then selects the "Trend of the Year" option.
[0635] 6. The server generates the video content based on the selected style and options.
[0636] 7. The user reviews the generated footage and requests corrections if necessary.
[0637] 8. The server completes the edited video.
[0638] 9. The user selects "Private" and sets the video content to be shared only with family members.
[0639] 10. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[0640] The system allows users to record important events in their lives and easily share them with others while respecting privacy.
[0641] The processing flow will be explained below.
[0642] Step 1:
[0643] A user visits the system's website or app and enters their login information.
[0644] Step 2:
[0645] The server authenticates the login information and verifies the user's profile information.
[0646] Step 3:
[0647] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[0648] Step 4:
[0649] The user enters text responses to the AI interview questions.
[0650] Step 5:
[0651] The server analyzes the user's answers and generates and displays the next question (e.g., "What is the duration of the event?").
[0652] Step 6:
[0653] The user answers additional questions as the interview continues.
[0654] Step 7:
[0655] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[0656] Step 8:
[0657] The server receives the uploaded material and stores it in storage.
[0658] Step 9:
[0659] The server analyzes the stored material and extracts the information necessary to generate the video.
[0660] Step 10:
[0661] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[0662] Step 11:
[0663] The user selects their preferred style and options (such as what was popular at the time).
[0664] Step 12:
[0665] The server stores the user's selections and prepares the video for generation.
[0666] Step 13:
[0667] The server uses AI algorithms to generate video content based on user input and provided materials.
[0668] Step 14:
[0669] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[0670] Step 15:
[0671] The user can check the footage on the preview page and enter correction requests if necessary.
[0672] Step 16:
[0673] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[0674] Step 17:
[0675] The user makes a final confirmation and sets the public range (private, limited, public).
[0676] Step 18:
[0677] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[0678] Example 1
[0679] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0680] In modern society, there are limited ways to record special events and important moments in video and share them with others. As a result, users face challenges in easily creating and sharing video content, requiring significant effort and technical knowledge. Furthermore, there is a lack of flexible ways to change the style and image of video content, making it difficult to meet the diverse needs of users. Furthermore, from the perspective of privacy protection, a lack of systems that allow users to easily set and manage the scope of disclosure has been pointed out.
[0681] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0682] In this invention, the server includes means for a user to input information about events or important events in his or her life, means for generating and presenting questions to the user and collecting the information input by the user, means for uploading multiple digital data provided by the user, means for analyzing and saving the uploaded digital data, means for a generation device to generate video content based on the collected information and analyzed data, means for a display device to display the generated video content, means for setting a disclosure range, and means for making the video content available to other users in accordance with the disclosure range. This enables users, without special knowledge, to easily visualize their important events, express them in various styles, and share them with others while protecting their privacy.
[0683] A "user" is someone who uses the system to visualize their own events or important events and share them with others.
[0684] "Server" refers to the central computer that controls and manages the entire system, and analyzes and stores information entered by users and uploaded material data.
[0685] "Terminal" means a device that allows a user to access the system, input information, and view video content.
[0686] A "generation device" is a device or software that uses AI algorithms to generate video content based on information collected from users and materials provided by them.
[0687] A "display device" is a device for displaying generated video content to a user.
[0688] The "means for generating questions" is a function for generating appropriate questions based on information input by the user and presenting them to the user.
[0689] "Digital data" refers to multimedia materials such as photographs, audio, and music provided by users.
[0690] "Means for uploading" refers to the functionality that allows users to transfer digital data to the system.
[0691] "Means for analyzing and storing" refers to the function for analyzing uploaded digital data and storing it within the system.
[0692] The "means for setting the disclosure range" is a function for selecting and setting the disclosure range of the generated video content.
[0693] The "means for publishing according to the public range" is a function for making video content accessible to other users based on the public range that has been set.
[0694] This invention relates to a system in which users input information about events or important events in their lives, and AI generates video content based on that information and shares it with other users. This system mainly involves information exchange and processing between a server, a terminal, and the user.
[0695] System configuration
[0696] 1. User login
[0697] To access the system, the terminal displays a login screen. The user enters their email address and password, which the terminal sends to the server. The server authenticates them by looking up their email address in a database (e.g., MongoDB or MySQL), and if authentication is successful, allows the user to proceed to the next step.
[0698] 2. Conducting AI interviews
[0699] The server retrieves the profile information of the user who successfully logged in and generates appropriate questions using a natural language processing model (e.g., GPT-4). The server sends a series of questions to the device, which displays them to the user. The user answers the questions, and the device sends the answers to the server.
[0700] 3. Provision of Materials
[0701] Users use their devices to upload digital data such as their photos, audio files, and songs. The devices send this data to a server, which then analyzes the data using image processing libraries (e.g., OpenCV) and audio analysis libraries (e.g., Librosa) and stores it in a storage system (e.g., Amazon S3).
[0702] 4. Select output format
[0703] The device displays an interface for selecting the style and image of the video content, allowing the user to choose styles such as anime, TV drama, American comics, etc. Additional options such as the trends of the time and historical background can also be selected. The device sends the selection results to the server, which stores them in a database.
[0704] 5. Image generation and confirmation
[0705] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials. The server temporarily saves the generated video and sends a URL to the device. The device displays a video preview screen, which the user can check. If the user sends a correction request as needed, the server accepts the correction request and regenerates the video.
[0706] 6. Setting the visibility and publishing
[0707] The device displays an interface for setting the visibility of the video content. The user can select from options such as private, limited, or public. The device sends the selection result to the server, which then manages access rights for the video content based on the selected visibility.
[0708] Specific examples
[0709] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0710] 1. The user logs in and the system starts the AI interview.
[0711] 2. The server sends the user a question such as, "Which event from your graduation ceremony was most memorable?", and the user answers with a specific episode.
[0712] 3. The user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and their favorite songs from their device.
[0713] 4. The server analyzes the uploaded material and suggests a video style (e.g., anime style).
[0714] 5. The user selects an anime style and chooses "Trend of the Year" as an additional option.
[0715] 6. The server generates the video content based on the selected style and options, using DeepArt or StyleGAN.
[0716] 7. The user checks the generated video on their device and requests corrections if necessary.
[0717] 8. The server completes the edited video.
[0718] 9. The user selects "Private" and sets the video content to be shared only with family members.
[0719] 10. The server publishes the video based on the user's settings and manages it so that only users within the specified range of access can access the video.
[0720] Example prompts to input to the generative AI model
[0721] A possible prompt for a generative AI model (e.g., GPT-4) might look something like this:
[0722] The user is talking about their college graduation ceremony. Use the following elements to generate animated video content:
[0723] Photo 1: A photo of me with my friends at the graduation ceremony.
[0724] Photo 2: Exterior of the ceremony venue.
[0725] Audio file: Audio of the speech on graduation day.
[0726] Song: A memorable song from graduation ceremony.
[0727] The tone of the video should be emotive and reflect the trends of the time.
[0728] This system allows users, without any special knowledge, to easily visualize their important events, express them in a variety of styles, and share them with others while protecting their privacy.
[0729] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0730] System program processing flow
[0731] Step 1:
[0732] User login
[0733] The device displays the login screen.
[0734] Input: User's email address and password.
[0735] Specific operation: Display a login form using HTML / CSS / JavaScript.
[0736] The user enters their email address and password.
[0737] Output: The login information entered.
[0738] The terminal sends the input information to the server.
[0739] Input: The login information entered.
[0740] Output: HTTP request to the server.
[0741] The server performs authentication by referencing a database (e.g. MongoDB or MySQL).
[0742] Input: The login information sent to the server.
[0743] Specific behavior: Provides an API (e.g., RESTful API) on the backend and handles authentication logic.
[0744] Output: Authentication result (success / failure).
[0745] The server returns the authentication result to the terminal, and if successful, proceeds to the next step.
[0746] Input: Authentication result.
[0747] Output: Next screen based on authentication result (instructions to proceed to next step if successful).
[0748] Step 2:
[0749] Conducting AI interviews
[0750] The server retrieves the user's profile information after a successful login.
[0751] Input: Login success notification.
[0752] Specific Actions: Retrieve user profile from database.
[0753] Output: Profile information.
[0754] The server uses the profile information to generate appropriate questions using a natural language processing model (e.g., GPT-4).
[0755] Input: Profile information.
[0756] What it does: Generate questions using GPT-4.
[0757] Output: Question list.
[0758] The server sends the generated question to the terminal.
[0759] Input: Questionnaire.
[0760] Output: HTTP response containing the questionnaire.
[0761] The device displays a question and the user enters the answer.
[0762] Input: Questionnaire.
[0763] Specific operation: Dynamically generate and display a question form.
[0764] Output: The user's answer.
[0765] The device sends the user's answer to the server.
[0766] Input: The user's answer.
[0767] Output: HTTP request to the server (response data).
[0768] Step 3:
[0769] Provision of materials
[0770] The device will display the interface for uploading materials.
[0771] Input: The upload request.
[0772] Specific behavior: HTML <input type="file"> Implement file selection using an element.
[0773] Output: Upload screen.
[0774] Users select and upload digital data such as photos, audio, and music.
[0775] Input: Materials (photos, audio, music, etc.).
[0776] Output: The digital data to be uploaded.
[0777] The device sends the uploaded data to the server.
[0778] Input: Digital data to be uploaded.
[0779] Output: HTTP request to server (digital data).
[0780] The server receives the data and stores it in a repository.
[0781] Input: Uploaded digital data.
[0782] Specific behavior: Stores data in a file storage system (e.g., Amazon S3).
[0783] Output: Stored digital data.
[0784] The server analyzes the stored data.
[0785] Input: Stored digital data.
[0786] Specific operation: Analyze the data using an image processing library (e.g., OpenCV) or a sound analysis library (e.g., Librosa).
[0787] Output: Analysis results.
[0788] Step 4:
[0789] Selecting the output format
[0790] The terminal displays an interface for selecting the style of the video content.
[0791] Input: A style selection request.
[0792] What it does: Provide users with choices using drop-down menus or radio buttons.
[0793] Output: Style selection screen.
[0794] The user selects a style and image from the options provided.
[0795] Enter: Style selection.
[0796] Output: Selected style information.
[0797] The terminal transmits the selection result to the server.
[0798] Input: Selected style information.
[0799] Output: HTTP request to the server (style information).
[0800] The server stores the user's selection in a database.
[0801] Input: Selected style information.
[0802] Specific action: Save to database.
[0803] Output: The saved style information.
[0804] Step 5:
[0805] Image generation and confirmation
[0806] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials.
[0807] Input: User input information and materials provided.
[0808] Specific operation: Generate images using DeepArt and StyleGAN.
[0809] Output: The generated video content.
[0810] The server temporarily stores the generated video and sends its URL to the terminal.
[0811] Input: Generated video content.
[0812] Output: URL of the archived video.
[0813] The device displays a video preview screen, and the user checks the video.
[0814] Input: The URL of the downloaded video.
[0815] Specific behavior: HTML <video>Display video using tags.
[0816] Output: The confirmation status of the user.
[0817] The user sends a modification request to the server as needed.
[0818] Input: The user's correction request.
[0819] Output: HTTP request to the server (modification request).
[0820] Step 6:
[0821] Setting the visibility and publishing
[0822] The device displays an interface for selecting the disclosure range.
[0823] Input: Disclosure scope setting request.
[0824] Specific behavior: Provide options for the scope of disclosure using radio buttons or checkboxes.
[0825] Output: Public range selection screen.
[0826] The user chooses the scope of disclosure.
[0827] Input: Select the disclosure range.
[0828] Output: Selected disclosure information.
[0829] The terminal transmits the selection result to the server.
[0830] Input: Selected disclosure information.
[0831] Output: HTTP request to the server (public range information).
[0832] The server sets the video content to be made public based on the selected public range.
[0833] Input: Selected disclosure information.
[0834] Specific operation: Set the scope of disclosure using ACL (Access Control List).
[0835] Output: The set disclosure range.
[0836] The server manages access rights for other users based on the public settings.
[0837] Input: The set disclosure range.
[0838] Output: Managed access rights.
[0839] (Application example 1)
[0840] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0841] In today's virtual stores, it is difficult to provide a personalized shopping experience based on users' emotions and personal memories. Therefore, a system that allows users to enjoy shopping while feeling an emotional connection is needed. Furthermore, existing systems lack the means to effectively utilize user-provided materials, display them as video content, and provide product suggestions and special offers. Furthermore, the technology to visualize individual user events and turning points and provide them as moving experiences within the virtual store is underdeveloped.
[0842] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0843] In this invention, the server includes: means for a user to input information about events and turning points that occurred in the user's life; means for a generation device to generate video content based on the input information; means for a display device to display the generated video content; means for setting a disclosure range; means for disclosing the video content to other users according to the disclosure range; and means for displaying products and special offers using the generated video content so that the user can personalize an emotional shopping experience in the virtual store. This allows the user to enjoy a virtual shopping experience based on personal memories while feeling an emotional connection, and further allows the user to effectively receive product suggestions and special offers through the video content.
[0844] "User" refers to an individual who uses the System.
[0845] "Events" refer to memorable experiences or turning points that occurred in the user's life.
[0846] A "turning point" refers to a time or event in a user's life when an important change or decision was made.
[0847] "Means for inputting information" refers to the interface through which a user provides information about their event to the system.
[0848] "Generation device" refers to a computer program and hardware system that creates video content based on input information.
[0849] "Display device" refers to a device that visually presents generated video content to a user.
[0850] The "means for setting the public range" refers to an interface that allows a user to select the viewing range of the generated video content.
[0851] "Publicity" refers to the range of users who can view the generated video content.
[0852] "Virtual store" refers to a virtual shopping environment provided on the Internet.
[0853] "Shopping Experience" refers to the series of activities in which a user browses, selects, and purchases products within a virtual store.
[0854] "Means for displaying products and special offers" refers to an interface for presenting related products and discount information within the generated video content.
[0855] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content and displays products and special offers in a virtual store. Specific embodiments of the system are described below.
[0856] System configuration
[0857] 1. User login
[0858] Users use the interface to log in to the system. When the user enters their login information, the server performs authentication.
[0859] 2. Conducting AI interviews
[0860] After logging in, the server checks the user's profile information and begins an AI interview. The server sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[0861] 3. Provision of Materials
[0862] Users upload digital data such as photos, audio, and favorite music through their devices, and the server stores and analyzes the uploaded material.
[0863] 4. Select output format
[0864] The server provides an interface for users to select the style and image of the video content (e.g., anime, TV drama, American comic book, etc.) Users can also select additional options such as the trends of the time and the historical background.
[0865] 5. Video Generation and Application to Virtual Stores
[0866] The server uses a generative AI model to generate video content based on user input and provided materials, which is then used to display products and special offers within the virtual store.
[0867] Hardware and software used
[0868] Hardware: smartphones, tablets, servers
[0869] Software: Online platform APIs, generative AI models (e.g., OpenAI GPT-3)
[0870] Data processing: Analyzes user input data (text, images, audio) and performs processes ranging from prompt generation to video content generation.
[0871] Specific examples
[0872] For example, consider the case where a user wants to create a video of their "college graduation ceremony." After logging in, the user answers questions such as "Which event from your graduation ceremony was most memorable?" through an AI interview. The server analyzes the materials provided by the user, such as photos, audio files, and favorite songs, and recommends video styles based on these. If the user selects an anime-style style and also selects "Trend of the Year" as an option, the server uses a generative AI model to generate video content based on the selected style and options. This video is then used to display products and special offers within the virtual store.
[0873] Prompt Sentence Examples
[0874] text
[0875] Title: College Graduation
[0876] Details: Wonderful days spent with many friends, and a graduation day that will remain in my memory forever.
[0877] Image files: ["graduation_photo1.jpg", "graduation_photo2.jpg"]
[0878] Audio file: ["graduation_speech.mp3"]
[0879] Style: Emotional documentary
[0880] Generate effective footage.
[0881] In this way, a system is realized that provides an emotional shopping experience in a virtual store based on the user's special events.
[0882] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0883] Step 1:
[0884] A user logs in to the system.
[0885] Input: User login information (username, password).
[0886] Output: If authentication is successful, proceed to next step. If not, prompt to retry.
[0887] Specific operation: The user enters the username and password into the login interface, and the server authenticates it. If the authentication is successful, the server obtains the user's profile information and proceeds to the next step.
[0888] Step 2:
[0889] AI interview begins.
[0890] Input: User profile information and login state.
[0891] Output: Detailed information about the user's events and turning points.
[0892] What it does: After logging in, the server asks the user a series of questions, which the user answers, and the server collects these answers, recording details of events and turning points that the user enters during this process.
[0893] Step 3:
[0894] Users upload materials.
[0895] Input: Digital data provided by the user, such as photos, audio, or favorite songs.
[0896] Output: Uploaded material is stored on the server and analyzed.
[0897] How it works: Users use their devices to upload photos, audio files, music, and other content to a server. The server then analyzes and stores the data. Image recognition and audio analysis algorithms are used for the analysis.
[0898] Step 4:
[0899] Select the style of your video content.
[0900] Input: User's event details and uploaded materials.
[0901] Output: User-selected video style and additional options.
[0902] Specific operation: Through the interface, the server allows the user to select the style of the video content (e.g., anime, drama, American comics, etc.). Additional options include the selection of the trends of the time and the historical background.
[0903] Step 5:
[0904] Video content generation.
[0905] Input: User event details, uploaded footage, selected video style and additional options.
[0906] Output: The generated video content.
[0907] Specific operation: The server uses the generative AI model to generate video content based on the user's input information and materials, inputting prompt sentences into the generative AI model and outputting video data based on them.
[0908] Step 6:
[0909] Displaying products and special offers in a virtual store.
[0910] Input: Generated video content.
[0911] Output: Products and special offers displayed in a virtual store with emotional video content.
[0912] What it does: The generated video content is played in the virtual store, where product suggestions and special offers related to the user's current situation are displayed, allowing the user to browse and purchase products with an emotional connection.
[0913] Step 7:
[0914] Setting the visibility of video content and publishing it.
[0915] Input: Generated video content, user-selected visibility.
[0916] Output: Video content with access rights controlled for other users based on the set visibility.
[0917] Specific operation: The server provides an interface for users to set the visibility of video content (private, limited, public, etc.). Based on the visibility set by the user, the server makes the video available to other users and manages access rights.
[0918] Through these steps, the system can visualize the user's special events and provide an emotional shopping experience within the virtual store.
[0919] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0920] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[0921] System configuration
[0922] 1. User login
[0923] A user visits the system's website or app and enters their login information.
[0924] 2. Conducting AI interviews
[0925] After logging in, the server checks the user's profile information and begins the AI interview.
[0926] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[0927] 3. Emotion recognition using the emotion engine
[0928] The server activates an emotion engine to recognize the user's emotion from the content of the user's answers and the provided materials.
[0929] The emotion engine analyzes the user's emotional state based on the context, words used, tone of voice, etc.
[0930] 4. Provision of Materials
[0931] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[0932] The server stores and analyzes the uploaded material.
[0933] 5. Select output format
[0934] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[0935] The server proposes an appropriate video style based on the results of the emotion engine.
[0936] 6. Image generation and confirmation
[0937] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[0938] Users can check the generated video on the preview page and request corrections if necessary.
[0939] 7. Setting the visibility and publishing
[0940] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[0941] The server makes the video content available to other users and manages access rights based on the settings.
[0942] 8. User response monitoring
[0943] The server monitors the user's reaction to the generated video content and collects emotional information about the user.
[0944] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[0945] Specific examples
[0946] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[0947] 1. The user logs in and the AI interview begins.
[0948] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[0949] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[0950] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[0951] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[0952] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[0953] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[0954] 8. The user reviews the generated footage and requests corrections if necessary.
[0955] 9. The server completes the edited video.
[0956] 10. The user selects "Private" and sets the video content to be shared only with family members.
[0957] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[0958] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[0959] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[0960] The processing flow will be explained below.
[0961] Step 1:
[0962] A user visits the system's website or app and enters their login information.
[0963] Step 2:
[0964] The server authenticates the login information and verifies the user's profile information.
[0965] Step 3:
[0966] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[0967] Step 4:
[0968] The user enters text responses to the AI interview questions.
[0969] Step 5:
[0970] The server receives the user's response and activates the emotion engine to analyze the user's emotional state. During this process, emotions are recognized from the content and context of the user's response, the words used, and the tone of voice.
[0971] Step 6:
[0972] The server stores the analysis results of the emotion engine and generates and displays the next question (e.g., "What is the duration of the event?").
[0973] Step 7:
[0974] The user answers additional questions as the interview continues.
[0975] Step 8:
[0976] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[0977] Step 9:
[0978] The server receives the uploaded material and stores it in storage.
[0979] Step 10:
[0980] The server analyzes the stored material and extracts the information necessary to generate the video.
[0981] Step 11:
[0982] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[0983] Step 12:
[0984] The user selects their preferred style and options (such as what was popular at the time).
[0985] Step 13:
[0986] The server presents a suggested output format (e.g., an emotional anime style) based on the analysis results of the emotion engine.
[0987] Step 14:
[0988] The server stores the user's selections and prepares the video for generation.
[0989] Step 15:
[0990] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[0991] Step 16:
[0992] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[0993] Step 17:
[0994] The user can check the footage on the preview page and enter correction requests if necessary.
[0995] Step 18:
[0996] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[0997] Step 19:
[0998] The user makes a final confirmation and sets the public range (private, limited, public).
[0999] Step 20:
[1000] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[1001] Step 21:
[1002] The server monitors viewers' reactions to the published video content and collects user emotional information.
[1003] Step 22:
[1004] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[1005] Example 2
[1006] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1007] Today's users have a growing need to digitize the events and turning points in their lives and save and share them as emotionally charged video content. However, existing systems lacked the technology to properly analyze users' emotions and reflect them in the video content. Furthermore, they lacked the means to generate different types of video content and allow users to easily set the visibility of their content, which hindered the user experience.
[1008] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input information about events and turning points that occurred in the user's life, a means for the server to generate video content using an AI algorithm based on the input information, and a means for an emotion engine to recognize emotions from the user's responses and provided materials. This enables a user to generate video content that reflects their own emotions, select a video style with a different image, easily set the scope of publication, and share it with other users.
[1009] A "user" is an individual or organization that uses the system and provides information about events or turning points that have occurred in their life.
[1010] "Input means" refers to the interface and devices that allow users to input information about events and turning points in their lives into the system.
[1011] "Generation means" refers to the function and device by which the server generates video content using an AI algorithm based on input information.
[1012] "Emotion Engine" refers to software and systems for recognizing and analyzing emotions from user responses and provided materials.
[1013] The "display means" refers to a device and interface for visually presenting the generated video content to the user.
[1014] The "publication range setting means" is an interface and device for allowing a user to set the publicity range of video content created by the user.
[1015] The "publication means" refers to a function and device that makes video content public to other users according to a set publicity range.
[1016] "Materials" is a general term for digital data such as photographs, audio data, and music provided by users.
[1017] "Uploading means" refers to the interface and device through which users send materials from their own terminals to the system.
[1018] "Analysis means" refers to a function and device that analyzes uploaded materials and stores them in a database.
[1019] The "image selection means" refers to an interface and device that allows the user to select different video styles (e.g., anime style, drama style, American comic style, etc.) when generating video content.
[1020] "Style change means" refers to functions and devices that automatically change the style of video content based on user selection.
[1021] The present invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, a server uses an AI algorithm to generate video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system is implemented as follows.
[1022] The system works through the website or application that the user accesses. The user first enters their email address and password on the login screen and is authenticated by the server. Once authenticated, the user is redirected to an interface where they can take an AI interview.
[1023] The server checks the user's profile information and uses a generative AI model (e.g., a general-purpose conversation model) to send the user a series of questions. When the user answers questions such as "What was the most memorable event during your time at university?", the server stores the answers in a database and passes them to an emotion engine. The emotion engine uses Google's emotion analysis API or similar to analyze the context and tone of the answers to identify the user's emotional state.
[1024] Next, users upload materials (digital data such as photos, audio files, and music) from their devices through the interface. The devices then send the materials to the server, which analyzes and stores them. Image recognition and audio analysis technologies are used for the analysis.
[1025] Next, the server suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred style from the presented styles. The selected style and options are saved on the server and used for video generation.
[1026] The server combines the selected style, emotional state, and user-provided materials to generate video content using a generative AI model (e.g., DeepMotion's AI animation generation technology). The user can view the generated video on a preview page and request corrections if necessary. The server applies the corrections and regenerates the video.
[1027] Finally, the server displays an interface for setting the visibility of the video content, and the user selects the visibility. The server then makes the video available to other users based on the settings and manages access rights. In this case, the server manages access rights for the video content according to the visibility, allowing for limited or general visibility.
[1028] In addition, the server monitors viewers' reactions to the video content after it has been released, reanalyzing the collected emotional information and saving it as reference information for the next video. This process allows users to create video content that reflects their own emotions and share it with other users within appropriate public limits.
[1029] As a concrete example, consider the case where a user wants to create a video of their college graduation ceremony. First, the user logs in, and the server sends a question such as, "Which event from your graduation ceremony was most memorable?" After the user answers with a specific episode, the server performs sentiment analysis to identify the user's emotions (e.g., joy, excitement, nostalgia). Next, the user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and favorite songs. The server analyzes these materials and recommends video styles based on the sentiment analysis results. The user selects a suggested style, and the server uses AI technology to generate video content based on the selected style and options. The user reviews the generated video and makes any necessary adjustments. Finally, the user selects "private" and shares the video content only with family members. The server manages the video based on the settings, monitors viewer reactions, and uses them to generate the next video. In this way, users can easily create and share emotionally relevant videos of their important events.
[1030] An example of a prompt sentence is asking the user, "Which event from your college graduation ceremony was the most memorable?" In this way, detailed information about the user's experience is collected, and emotion analysis and video generation are performed based on that information.
[1031] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1032] Step 1:
[1033] A user accesses the system's website or app and enters their email address and password on the login screen. The server compares this information with the database, and if authentication is successful, the user is taken to the dashboard screen.
[1034] Input: User's email address and password
[1035] Data processing: Check the input information against the database
[1036] Output: User authentication result, dashboard screen displayed if authentication is successful
[1037] Step 2:
[1038] After logging in, the server checks the user's profile information and uses the generative AI model to initiate an AI interview, sending the user a series of questions to which they respond.
[1039] Input: User profile information, AI model prompt
[1040] Data processing: Generate appropriate questions based on AI algorithms
[1041] Output: Sending interview questions to the user
[1042] Step 3:
[1043] The user answers questions sent by the server. The answers are entered as text or voice. The server receives these answers and stores them in a database.
[1044] Input: User response (text or voice)
[1045] Data processing: Saving response data
[1046] Output: Response data stored in a database
[1047] Step 4:
[1048] The server passes the user's response data to the emotion engine, which uses natural language processing technology to analyze the context and tone of the response and identify the emotional state (e.g., joy, sadness, surprise).
[1049] Input: User response data
[1050] Data processing: Sentiment analysis (using NLP technology)
[1051] Output: Identified emotional state data
[1052] Step 5:
[1053] Users upload digital materials such as photos, audio files, and music from their own devices through the interface. The devices send this data to the server, which then analyzes and stores it in a database.
[1054] Input: User photos, audio files, songs
[1055] Data processing: Analysis and storage of material data
[1056] Output: Material data stored in a database
[1057] Step 6:
[1058] The server displays an interface that suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred video style from the presented options.
[1059] Input: User sentiment analysis results, existing video style data
[1060] Data processing: Proposing appropriate video styles
[1061] Output: User suggestions, user choices
[1062] Step 7:
[1063] The server integrates the selected style, emotional state, and user-provided materials to generate video content using a generative AI model, and the user can view the generated video content on a preview page.
[1064] Input: Selected visual style, emotional state, user material data
[1065] Data processing: Image generation using AI algorithms
[1066] Output: Generated video content
[1067] Step 8:
[1068] The user checks the generated video content and requests corrections if necessary. The server receives the correction request and regenerates the video.
[1069] Input: User's correction request
[1070] Data processing: Regenerating a modified version of the video
[1071] Output: Modified video content
[1072] Step 9:
[1073] The server displays an interface for setting the visibility of video content, and the user selects a setting from options such as private, limited, or public. Based on the setting, the video is made public to other users and access rights are managed.
[1074] Input: User's visibility settings
[1075] Data processing: Applying disclosure settings
[1076] Output: Published video content
[1077] Step 10:
[1078] After the video content is released, the server monitors the viewer's reaction to it, reanalyzes it using the emotion engine, and saves the information as reference for the next video generation.
[1079] Input: Viewer response data (likes, comments, etc.)
[1080] Data processing: sentiment analysis, data storage
[1081] Output: Reference data for future video generation
[1082] In this way, users can create and share video content that reflects their own emotions.
[1083] (Application example 2)
[1084] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1085] In today's world, people have a growing desire to record and share important events and turning points in their lives as video. However, this process requires a lot of time and effort, and it is particularly challenging to generate video that expresses appropriate emotions for each event. Furthermore, determining the extent to which the generated video should be shared and monitoring viewer reactions to collect emotions are also complex. A simple and effective method to solve these problems is needed.
[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input information about events and turning points that occurred in the user's life, means for a generation device to generate video content based on the input information, means for a display device to display the generated video content, means for setting a disclosure range, means for disclosing the video content to other users according to the disclosure range, and means for monitoring viewer reactions and collecting user emotional information. This enables a user to easily generate video content that reflects the emotions of important events, share it within an appropriate disclosure range, and further collect viewer reactions to use in generating the next video.
[1087] "User" refers to an individual or group that uses the system to input information about events and turning points that have occurred in their life, and to generate and share video content.
[1088] "Generation device" refers to a device or software that automatically generates video content based on information input by a user.
[1089] A "display device" is a device or interface for visually displaying generated video content to a user.
[1090] "Means for setting the scope of disclosure" refers to a function or interface that allows users to select and set to whom and to what extent the generated video content will be disclosed.
[1091] "Viewers" refer to other users (third parties) who view published video content.
[1092] "Monitoring" refers to the process of observing and recording audience reactions and behavior, especially capturing audience emotional responses.
[1093] "Emotional information" refers to data or analytical results that indicate the emotional state of a user or viewer.
[1094] "Multiple Materials" refers to digital media files such as photos, audio, and music uploaded by users.
[1095] A "generative AI model" is an artificial intelligence technology that generates a video script based on information input by the user and emotional analysis.
[1096] A "prompt sentence" is an input text given to a generative AI model to instruct it on the content and style of the video content to be generated.
[1097] This invention relates to a system in which users input information about events and turning points in their lives, and AI generates video content based on that information, and an emotion engine is used to recognize and share the user's emotions. This system involves a series of processes, including user information input, preview of the generated video, setting the sharing range, and monitoring viewer reactions.
[1098] System configuration
[1099] User login
[1100] The server provides a means for users to access the system's website or application and enter their login information, which receives the user's authentication information and logs them into the system.
[1101] Conducting AI interviews
[1102] After logging in, the server checks the user's profile information and begins an AI interview. Specifically, it sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[1103] Emotion recognition with emotion engine
[1104] The server then activates an emotion engine based on the user's responses to recognize the user's emotions. The emotion engine analyzes the user's emotional state from their words and tone of voice.
[1105] Provision of materials and analysis
[1106] Users upload materials for video content, such as photos, audio, and music, from their devices. The server stores and analyzes the uploaded materials.
[1107] Selecting the output format
[1108] Users can use an interface to select the style of the video content (e.g., anime, drama, comic, etc.). The server then suggests an appropriate video style based on the results of the emotion engine.
[1109] Image generation and confirmation
[1110] The server uses a generative AI model to generate video content based on user input, provided materials, and recognized emotions. Users can view the generated video on a preview page and request corrections if necessary.
[1111] Setting the visibility and publishing
[1112] The server displays an interface for setting the visibility of the video content, and the user can select the visibility (e.g., private, limited, public). The server makes the video content available to other users based on the settings and manages access rights.
[1113] User response monitoring
[1114] The server monitors the viewer's reaction to the generated video content and collects information on the user's emotions, which is then saved as reference for the next video generation.
[1115] Hardware and software used
[1116] The system is implemented using the following hardware and software:
[1117] Servers: Web servers, database servers
[1118] Device: PC, smartphone, tablet, etc. that users access
[1119] Software: Generative AI models (e.g., OpenAI), emotion analysis engines (e.g., emotion_recognition library)
[1120] Specific examples
[1121] For example, a user may want to film their "college graduation ceremony."
[1122] procedure
[1123] 1. The user logs in and the AI interview begins.
[1124] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[1125] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[1126] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[1127] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[1128] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[1129] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[1130] 8. The user reviews the generated footage and requests corrections if necessary.
[1131] 9. The server completes the edited video.
[1132] 10. The user selects "Private" and sets the video content to be shared only with family members.
[1133] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[1134] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[1135] Prompt Sentence Examples
[1136] For example, the following prompt sentence is input to the generative AI model:
[1137] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[1138] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[1139] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1140] Step 1:
[1141] A user visits the system's website or application and enters their login information. The input to this step is the user's credentials (e.g., email address and password), and the output is a message indicating authentication success or failure. The server checks the user's credentials against information in its database and, if there is a match, allows the user to log in.
[1142] Step 2:
[1143] After logging in, the server checks the user's profile information and starts the AI interview. The input is the user's profile information and a pre-registered question list, and the output is sending questions to the user. The server sequentially sends the user questions such as "Which event at your graduation ceremony was most memorable?"
[1144] Step 3:
[1145] The user answers questions from the server. The input of this step is the user's text answer, and the output is the collected episode details. The user answers specific episodes based on their own experiences, and the information is sent to the server.
[1146] Step 4:
[1147] The server activates an emotion engine based on the collected responses to recognize the user's emotions. The input of this step is the user's response information, and the output is analyzed emotion data. The emotion engine analyzes the language tone and keywords from the user's response to identify the user's emotional state.
[1148] Step 5:
[1149] Users upload materials for video content, such as photos, audio, and music, from their devices to the server. The input is the digital media files (photos, audio, and music) uploaded by the user, and the output is a message confirming the completion of the upload. The device sends the provided materials to the server, which then stores them.
[1150] Step 6:
[1151] The server analyzes the uploaded material and proposes a style for the video content based on it. The input of this step is the uploaded material and emotional data, and the output is a recommended video style. The server analyzes the user's material and emotional state and recommends an appropriate style (e.g., an emotional anime style).
[1152] Step 7:
[1153] The user selects a suggested video style and then selects further options. The input is the style suggestion from the server and the user's selection, and the output is the selected style and options. The user makes their selection through the interface, and that information is sent to the server.
[1154] Step 8:
[1155] The server generates video content using a generative AI model based on the selected style and options. The input for this step is the user's selection information and emotional data, which are converted into specific prompt sentences. The output is the generated video content. The server sends the following prompt sentence to the generative AI model:
[1156] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[1157] Step 9:
[1158] The user can check the generated video on the preview page and request corrections if necessary. The input is the generated video content, and the output is the user's feedback or correction requests. The server displays the generated video to the user and receives any corrections that need to be made.
[1159] Step 10:
[1160] The server completes the revised video. The input of this step is the user's feedback, and the output is the completed video content. The server regenerates the video based on the user's feedback and provides the final version.
[1161] Step 11:
[1162] The user sets the visibility of the completed video content. The input is the visibility selection, and the output is a confirmation message that the settings have been completed. The server centrally makes the video available to other users based on the visibility (private, limited, public) entered.
[1163] Step 12:
[1164] The server monitors viewers' reactions to video content and collects emotional information. The input is viewer reaction data, and the output is analyzed emotional information. The server monitors viewers' reactions and comments and records them as reference information for the next video generation.
[1165] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1166] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1167] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1168] [Third embodiment]
[1169] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1170] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1171] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1172] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1173] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1174] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1175] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1176] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1177] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1178] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1179] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1180] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1181] This invention relates to a system in which users input information about events and turning points that occurred in their lives, and AI generates video content based on that information and shares it with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[1182] System configuration
[1183] 1. User login
[1184] The user is presented with an interface to log into the system, where they enter their login information and are authenticated.
[1185] 2. Conducting AI interviews
[1186] After logging in, the server checks the user's profile information and begins the AI interview.
[1187] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[1188] 3. Provision of Materials
[1189] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[1190] The server stores and analyzes the uploaded material.
[1191] 4. Select output format
[1192] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[1193] Users can also select additional options such as trends and historical background of the time.
[1194] 5. Image generation and confirmation
[1195] The server uses AI algorithms to generate video content based on user input and provided materials.
[1196] Users can check the generated video on the preview page and request corrections if necessary.
[1197] 6. Setting the visibility and publishing
[1198] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[1199] The server makes the video content available to other users and manages access rights based on the settings.
[1200] Specific examples
[1201] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[1202] 1. The user logs in and the AI interview begins.
[1203] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[1204] 3. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[1205] 4. The server analyzes the user's material and recommends a video style (for example, anime style) based on that analysis.
[1206] 5. The user selects an anime style and then selects the "Trend of the Year" option.
[1207] 6. The server generates the video content based on the selected style and options.
[1208] 7. The user reviews the generated footage and requests corrections if necessary.
[1209] 8. The server completes the edited video.
[1210] 9. The user selects "Private" and sets the video content to be shared only with family members.
[1211] 10. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[1212] The system allows users to record important events in their lives and easily share them with others while respecting privacy.
[1213] The processing flow will be explained below.
[1214] Step 1:
[1215] A user visits the system's website or app and enters their login information.
[1216] Step 2:
[1217] The server authenticates the login information and verifies the user's profile information.
[1218] Step 3:
[1219] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[1220] Step 4:
[1221] The user enters text responses to the AI interview questions.
[1222] Step 5:
[1223] The server analyzes the user's answers and generates and displays the next question (e.g., "What is the duration of the event?").
[1224] Step 6:
[1225] The user answers additional questions as the interview continues.
[1226] Step 7:
[1227] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[1228] Step 8:
[1229] The server receives the uploaded material and stores it in storage.
[1230] Step 9:
[1231] The server analyzes the stored material and extracts the information necessary to generate the video.
[1232] Step 10:
[1233] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[1234] Step 11:
[1235] The user selects their preferred style and options (such as what was popular at the time).
[1236] Step 12:
[1237] The server stores the user's selections and prepares the video for generation.
[1238] Step 13:
[1239] The server uses AI algorithms to generate video content based on user input and provided materials.
[1240] Step 14:
[1241] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[1242] Step 15:
[1243] The user can check the footage on the preview page and enter correction requests if necessary.
[1244] Step 16:
[1245] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[1246] Step 17:
[1247] The user makes a final confirmation and sets the public range (private, limited, public).
[1248] Step 18:
[1249] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[1250] Example 1
[1251] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1252] In modern society, there are limited ways to record special events and important moments in video and share them with others. As a result, users face challenges in easily creating and sharing video content, requiring significant effort and technical knowledge. Furthermore, there is a lack of flexible ways to change the style and image of video content, making it difficult to meet the diverse needs of users. Furthermore, from the perspective of privacy protection, a lack of systems that allow users to easily set and manage the scope of disclosure has been pointed out.
[1253] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1254] In this invention, the server includes means for a user to input information about events or important events in his or her life, means for generating and presenting questions to the user and collecting the information input by the user, means for uploading multiple digital data provided by the user, means for analyzing and saving the uploaded digital data, means for a generation device to generate video content based on the collected information and analyzed data, means for a display device to display the generated video content, means for setting a disclosure range, and means for making the video content available to other users in accordance with the disclosure range. This enables users, without special knowledge, to easily visualize their important events, express them in various styles, and share them with others while protecting their privacy.
[1255] A "user" is someone who uses the system to visualize their own events or important events and share them with others.
[1256] "Server" refers to the central computer that controls and manages the entire system, and analyzes and stores information entered by users and uploaded material data.
[1257] "Terminal" means a device that allows a user to access the system, input information, and view video content.
[1258] A "generation device" is a device or software that uses AI algorithms to generate video content based on information collected from users and materials provided by them.
[1259] A "display device" is a device for displaying generated video content to a user.
[1260] The "means for generating questions" is a function for generating appropriate questions based on information input by the user and presenting them to the user.
[1261] "Digital data" refers to multimedia materials such as photographs, audio, and music provided by users.
[1262] "Means for uploading" refers to the functionality that allows users to transfer digital data to the system.
[1263] "Means for analyzing and storing" refers to the function for analyzing uploaded digital data and storing it within the system.
[1264] The "means for setting the disclosure range" is a function for selecting and setting the disclosure range of the generated video content.
[1265] The "means for publishing according to the public range" is a function for making video content accessible to other users based on the public range that has been set.
[1266] This invention relates to a system in which users input information about events or important events in their lives, and AI generates video content based on that information and shares it with other users. This system mainly involves information exchange and processing between a server, a terminal, and the user.
[1267] System configuration
[1268] 1. User login
[1269] To access the system, the terminal displays a login screen. The user enters their email address and password, which the terminal sends to the server. The server authenticates them by looking up their email address in a database (e.g., MongoDB or MySQL), and if authentication is successful, allows the user to proceed to the next step.
[1270] 2. Conducting AI interviews
[1271] The server retrieves the profile information of the user who successfully logged in and generates appropriate questions using a natural language processing model (e.g., GPT-4). The server sends a series of questions to the device, which displays them to the user. The user answers the questions, and the device sends the answers to the server.
[1272] 3. Provision of Materials
[1273] Users use their devices to upload digital data such as their photos, audio files, and songs. The devices send this data to a server, which then analyzes the data using image processing libraries (e.g., OpenCV) and audio analysis libraries (e.g., Librosa) and stores it in a storage system (e.g., Amazon S3).
[1274] 4. Select output format
[1275] The device displays an interface for selecting the style and image of the video content, allowing the user to choose styles such as anime, TV drama, American comics, etc. Additional options such as the trends of the time and historical background can also be selected. The device sends the selection results to the server, which stores them in a database.
[1276] 5. Image generation and confirmation
[1277] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials. The server temporarily saves the generated video and sends a URL to the device. The device displays a video preview screen, which the user can check. If the user sends a correction request as needed, the server accepts the correction request and regenerates the video.
[1278] 6. Setting the visibility and publishing
[1279] The device displays an interface for setting the visibility of the video content. The user can select from options such as private, limited, or public. The device sends the selection result to the server, which then manages access rights for the video content based on the selected visibility.
[1280] Specific examples
[1281] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[1282] 1. The user logs in and the system starts the AI interview.
[1283] 2. The server sends the user a question such as, "Which event from your graduation ceremony was most memorable?", and the user answers with a specific episode.
[1284] 3. The user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and their favorite songs from their device.
[1285] 4. The server analyzes the uploaded material and suggests a video style (e.g., anime style).
[1286] 5. The user selects an anime style and chooses "Trend of the Year" as an additional option.
[1287] 6. The server generates the video content based on the selected style and options, using DeepArt or StyleGAN.
[1288] 7. The user checks the generated video on their device and requests corrections if necessary.
[1289] 8. The server completes the edited video.
[1290] 9. The user selects "Private" and sets the video content to be shared only with family members.
[1291] 10. The server publishes the video based on the user's settings and manages it so that only users within the specified range of access can access the video.
[1292] Example prompts to input to the generative AI model
[1293] A possible prompt for a generative AI model (e.g., GPT-4) might look something like this:
[1294] The user is talking about their college graduation ceremony. Use the following elements to generate animated video content:
[1295] Photo 1: A photo of me with my friends at the graduation ceremony.
[1296] Photo 2: Exterior of the ceremony venue.
[1297] Audio file: Audio of the speech on graduation day.
[1298] Song: A memorable song from graduation ceremony.
[1299] The tone of the video should be emotive and reflect the trends of the time.
[1300] This system allows users, without any special knowledge, to easily visualize their important events, express them in a variety of styles, and share them with others while protecting their privacy.
[1301] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1302] System program processing flow
[1303] Step 1:
[1304] User login
[1305] The device displays the login screen.
[1306] Input: User's email address and password.
[1307] Specific operation: Display a login form using HTML / CSS / JavaScript.
[1308] The user enters their email address and password.
[1309] Output: The login information entered.
[1310] The terminal sends the input information to the server.
[1311] Input: The login information entered.
[1312] Output: HTTP request to the server.
[1313] The server performs authentication by referencing a database (e.g. MongoDB or MySQL).
[1314] Input: The login information sent to the server.
[1315] Specific behavior: Provides an API (e.g., RESTful API) on the backend and handles authentication logic.
[1316] Output: Authentication result (success / failure).
[1317] The server returns the authentication result to the terminal, and if successful, proceeds to the next step.
[1318] Input: Authentication result.
[1319] Output: Next screen based on authentication result (instructions to proceed to next step if successful).
[1320] Step 2:
[1321] Conducting AI interviews
[1322] The server retrieves the user's profile information after a successful login.
[1323] Input: Login success notification.
[1324] Specific Actions: Retrieve user profile from database.
[1325] Output: Profile information.
[1326] The server uses the profile information to generate appropriate questions using a natural language processing model (e.g., GPT-4).
[1327] Input: Profile information.
[1328] What it does: Generate questions using GPT-4.
[1329] Output: Question list.
[1330] The server sends the generated question to the terminal.
[1331] Input: Questionnaire.
[1332] Output: HTTP response containing the questionnaire.
[1333] The device displays a question and the user enters the answer.
[1334] Input: Questionnaire.
[1335] Specific operation: Dynamically generate and display a question form.
[1336] Output: The user's answer.
[1337] The device sends the user's answer to the server.
[1338] Input: The user's answer.
[1339] Output: HTTP request to the server (response data).
[1340] Step 3:
[1341] Provision of materials
[1342] The device will display the interface for uploading materials.
[1343] Input: The upload request.
[1344] Specific behavior: HTML <input type="file"> Implement file selection using an element.
[1345] Output: Upload screen.
[1346] Users select and upload digital data such as photos, audio, and music.
[1347] Input: Materials (photos, audio, music, etc.).
[1348] Output: The digital data to be uploaded.
[1349] The device sends the uploaded data to the server.
[1350] Input: Digital data to be uploaded.
[1351] Output: HTTP request to server (digital data).
[1352] The server receives the data and stores it in a repository.
[1353] Input: Uploaded digital data.
[1354] Specific behavior: Stores data in a file storage system (e.g., Amazon S3).
[1355] Output: Stored digital data.
[1356] The server analyzes the stored data.
[1357] Input: Stored digital data.
[1358] Specific operation: Analyze the data using an image processing library (e.g., OpenCV) or a sound analysis library (e.g., Librosa).
[1359] Output: Analysis results.
[1360] Step 4:
[1361] Selecting the output format
[1362] The terminal displays an interface for selecting the style of the video content.
[1363] Input: A style selection request.
[1364] What it does: Provide users with choices using drop-down menus or radio buttons.
[1365] Output: Style selection screen.
[1366] The user selects a style and image from the options provided.
[1367] Enter: Style selection.
[1368] Output: Selected style information.
[1369] The terminal transmits the selection result to the server.
[1370] Input: Selected style information.
[1371] Output: HTTP request to the server (style information).
[1372] The server stores the user's selection in a database.
[1373] Input: Selected style information.
[1374] Specific action: Save to database.
[1375] Output: The saved style information.
[1376] Step 5:
[1377] Image generation and confirmation
[1378] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials.
[1379] Input: User input information and materials provided.
[1380] Specific operation: Generate images using DeepArt and StyleGAN.
[1381] Output: The generated video content.
[1382] The server temporarily stores the generated video and sends its URL to the terminal.
[1383] Input: Generated video content.
[1384] Output: URL of the archived video.
[1385] The device displays a video preview screen, and the user checks the video.
[1386] Input: The URL of the downloaded video.
[1387] Specific behavior: HTML <video>Display video using tags.
[1388] Output: The confirmation status of the user.
[1389] The user sends a modification request to the server as needed.
[1390] Input: The user's correction request.
[1391] Output: HTTP request to the server (modification request).
[1392] Step 6:
[1393] Setting the visibility and publishing
[1394] The device displays an interface for selecting the disclosure range.
[1395] Input: Disclosure scope setting request.
[1396] Specific behavior: Provide options for the scope of disclosure using radio buttons or checkboxes.
[1397] Output: Public range selection screen.
[1398] The user chooses the scope of disclosure.
[1399] Input: Select the disclosure range.
[1400] Output: Selected disclosure information.
[1401] The terminal transmits the selection result to the server.
[1402] Input: Selected disclosure information.
[1403] Output: HTTP request to the server (public range information).
[1404] The server sets the video content to be made public based on the selected public range.
[1405] Input: Selected disclosure information.
[1406] Specific operation: Set the scope of disclosure using ACL (Access Control List).
[1407] Output: The set disclosure range.
[1408] The server manages access rights for other users based on the public settings.
[1409] Input: The set disclosure range.
[1410] Output: Managed access rights.
[1411] (Application example 1)
[1412] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1413] In today's virtual stores, it is difficult to provide a personalized shopping experience based on users' emotions and personal memories. Therefore, a system that allows users to enjoy shopping while feeling an emotional connection is needed. Furthermore, existing systems lack the means to effectively utilize user-provided materials, display them as video content, and make product suggestions and special offers. Furthermore, the technology to visualize individual user events and turning points and provide them as moving experiences within the virtual store is underdeveloped.
[1414] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1415] In this invention, the server includes: means for a user to input information about events and turning points that occurred in the user's life; means for a generation device to generate video content based on the input information; means for a display device to display the generated video content; means for setting a disclosure range; means for disclosing the video content to other users according to the disclosure range; and means for displaying products and special offers using the generated video content so that the user can personalize an emotional shopping experience in the virtual store. This allows the user to enjoy a virtual shopping experience based on personal memories while feeling an emotional connection, and further allows the user to effectively receive product suggestions and special offers through the video content.
[1416] "User" refers to an individual who uses the System.
[1417] "Events" refer to memorable experiences or turning points that occurred in the user's life.
[1418] A "turning point" refers to a time or event in a user's life when an important change or decision was made.
[1419] "Means for inputting information" refers to the interface through which a user provides information about their event to the system.
[1420] "Generation device" refers to a computer program and hardware system that creates video content based on input information.
[1421] "Display device" refers to a device that visually presents generated video content to a user.
[1422] The "means for setting the public range" refers to an interface that allows a user to select the viewing range of the generated video content.
[1423] "Publicity" refers to the range of users who can view the generated video content.
[1424] "Virtual store" refers to a virtual shopping environment provided on the Internet.
[1425] "Shopping Experience" refers to the series of activities in which a user browses, selects, and purchases products within a virtual store.
[1426] "Means for displaying products and special offers" refers to an interface for presenting related products and discount information within the generated video content.
[1427] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content and displays products and special offers in a virtual store. Specific embodiments of the system are described below.
[1428] System configuration
[1429] 1. User login
[1430] Users use the interface to log in to the system. When the user enters their login information, the server performs authentication.
[1431] 2. Conducting AI interviews
[1432] After logging in, the server checks the user's profile information and begins an AI interview. The server sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[1433] 3. Provision of Materials
[1434] Users upload digital data such as photos, audio, and favorite music through their devices, and the server stores and analyzes the uploaded material.
[1435] 4. Select output format
[1436] The server provides an interface for users to select the style and image of the video content (e.g., anime, TV drama, American comic book, etc.) Users can also select additional options such as the trends of the time and the historical background.
[1437] 5. Video Generation and Application to Virtual Stores
[1438] The server uses a generative AI model to generate video content based on user input and provided materials, which is then used to display products and special offers within the virtual store.
[1439] Hardware and software used
[1440] Hardware: smartphones, tablets, servers
[1441] Software: Online platform APIs, generative AI models (e.g., OpenAI GPT-3)
[1442] Data processing: Analyzes user input data (text, images, audio) and performs processes ranging from prompt generation to video content generation.
[1443] Specific examples
[1444] For example, consider the case where a user wants to create a video of their "college graduation ceremony." After logging in, the user answers questions such as "Which event from your graduation ceremony was most memorable?" through an AI interview. The server analyzes the materials provided by the user, such as photos, audio files, and favorite songs, and recommends video styles based on these. If the user selects an anime-style style and also selects "Trend of the Year" as an option, the server uses a generative AI model to generate video content based on the selected style and options. This video is then used to display products and special offers within the virtual store.
[1445] Prompt Sentence Examples
[1446] text
[1447] Title: College Graduation
[1448] Details: Wonderful days spent with many friends, and a graduation day that will remain in my memory forever.
[1449] Image files: ["graduation_photo1.jpg", "graduation_photo2.jpg"]
[1450] Audio file: ["graduation_speech.mp3"]
[1451] Style: Emotional documentary
[1452] Generate effective footage.
[1453] In this way, a system is realized that provides an emotional shopping experience in a virtual store based on the user's special events.
[1454] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1455] Step 1:
[1456] A user logs in to the system.
[1457] Input: User login information (username, password).
[1458] Output: If authentication is successful, proceed to next step. If not, prompt to retry.
[1459] Specific operation: The user enters the username and password into the login interface, and the server authenticates it. If the authentication is successful, the server obtains the user's profile information and proceeds to the next step.
[1460] Step 2:
[1461] AI interview begins.
[1462] Input: User profile information and login state.
[1463] Output: Detailed information about the user's events and turning points.
[1464] What it does: After logging in, the server asks the user a series of questions, which the user answers, and the server collects these answers, recording details of events and turning points that the user enters during this process.
[1465] Step 3:
[1466] Users upload materials.
[1467] Input: Digital data provided by the user, such as photos, audio, or favorite songs.
[1468] Output: Uploaded material is stored on the server and analyzed.
[1469] How it works: Users use their devices to upload photos, audio files, music, and other content to a server. The server then analyzes and stores the data. Image recognition and audio analysis algorithms are used for the analysis.
[1470] Step 4:
[1471] Select the style of your video content.
[1472] Input: User's event details and uploaded materials.
[1473] Output: User-selected video style and additional options.
[1474] Specific operation: Through the interface, the server allows the user to select the style of the video content (e.g., anime, drama, American comics, etc.). Additional options include the selection of the trends of the time and the historical background.
[1475] Step 5:
[1476] Video content generation.
[1477] Input: User event details, uploaded footage, selected video style and additional options.
[1478] Output: The generated video content.
[1479] Specific operation: The server uses the generative AI model to generate video content based on the user's input information and materials, inputting prompt sentences into the generative AI model and outputting video data based on them.
[1480] Step 6:
[1481] Displaying products and special offers in a virtual store.
[1482] Input: Generated video content.
[1483] Output: Products and special offers displayed in a virtual store with emotional video content.
[1484] What it does: The generated video content is played in the virtual store, where product suggestions and special offers related to the user's current situation are displayed, allowing the user to browse and purchase products with an emotional connection.
[1485] Step 7:
[1486] Setting the visibility of video content and publishing it.
[1487] Input: Generated video content, user-selected visibility.
[1488] Output: Video content with access rights controlled for other users based on the set visibility.
[1489] Specific operation: The server provides an interface for users to set the visibility of video content (private, limited, public, etc.). Based on the visibility set by the user, the server makes the video available to other users and manages access rights.
[1490] Through these steps, the system can visualize the user's special events and provide an emotional shopping experience within the virtual store.
[1491] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1492] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[1493] System configuration
[1494] 1. User login
[1495] A user visits the system's website or app and enters their login information.
[1496] 2. Conducting AI interviews
[1497] After logging in, the server checks the user's profile information and begins the AI interview.
[1498] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[1499] 3. Emotion recognition using the emotion engine
[1500] The server activates an emotion engine to recognize the user's emotion from the content of the user's answers and the provided materials.
[1501] The emotion engine analyzes the user's emotional state based on the context, words used, tone of voice, etc.
[1502] 4. Provision of Materials
[1503] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[1504] The server stores and analyzes the uploaded material.
[1505] 5. Select output format
[1506] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[1507] The server proposes an appropriate video style based on the results of the emotion engine.
[1508] 6. Image generation and confirmation
[1509] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[1510] Users can check the generated video on the preview page and request corrections if necessary.
[1511] 7. Setting the visibility and publishing
[1512] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[1513] The server makes the video content available to other users and manages access rights based on the settings.
[1514] 8. User response monitoring
[1515] The server monitors the user's reaction to the generated video content and collects emotional information about the user.
[1516] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[1517] Specific examples
[1518] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[1519] 1. The user logs in and the AI interview begins.
[1520] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[1521] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[1522] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[1523] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[1524] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[1525] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[1526] 8. The user reviews the generated footage and requests corrections if necessary.
[1527] 9. The server completes the edited video.
[1528] 10. The user selects "Private" and sets the video content to be shared only with family members.
[1529] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[1530] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[1531] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[1532] The processing flow will be explained below.
[1533] Step 1:
[1534] A user visits the system's website or app and enters their login information.
[1535] Step 2:
[1536] The server authenticates the login information and verifies the user's profile information.
[1537] Step 3:
[1538] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[1539] Step 4:
[1540] The user enters text responses to the AI interview questions.
[1541] Step 5:
[1542] The server receives the user's response and activates the emotion engine to analyze the user's emotional state. During this process, emotions are recognized from the content and context of the user's response, the words used, and the tone of voice.
[1543] Step 6:
[1544] The server stores the analysis results of the emotion engine and generates and displays the next question (e.g., "What is the duration of the event?").
[1545] Step 7:
[1546] The user answers additional questions as the interview continues.
[1547] Step 8:
[1548] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[1549] Step 9:
[1550] The server receives the uploaded material and stores it in storage.
[1551] Step 10:
[1552] The server analyzes the stored material and extracts the information necessary to generate the video.
[1553] Step 11:
[1554] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[1555] Step 12:
[1556] The user selects their preferred style and options (such as what was popular at the time).
[1557] Step 13:
[1558] The server presents a suggested output format (e.g., an emotional anime style) based on the analysis results of the emotion engine.
[1559] Step 14:
[1560] The server stores the user's selections and prepares the video for generation.
[1561] Step 15:
[1562] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[1563] Step 16:
[1564] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[1565] Step 17:
[1566] The user can check the footage on the preview page and enter correction requests if necessary.
[1567] Step 18:
[1568] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[1569] Step 19:
[1570] The user makes a final confirmation and sets the public range (private, limited, public).
[1571] Step 20:
[1572] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[1573] Step 21:
[1574] The server monitors viewers' reactions to the published video content and collects user emotional information.
[1575] Step 22:
[1576] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[1577] Example 2
[1578] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1579] Today's users have a growing need to digitize the events and turning points in their lives and save and share them as emotionally charged video content. However, existing systems for achieving this lacked the technology to properly analyze users' emotions and reflect them in the video content. Furthermore, they lacked a means to generate video content with different images and allow users to easily set the visibility of the content, which hindered the user experience.
[1580] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input information about events and turning points that occurred in the user's life, a means for the server to generate video content using an AI algorithm based on the input information, and a means for an emotion engine to recognize emotions from the user's responses and provided materials. This enables a user to generate video content that reflects their own emotions, select a video style with a different image, easily set the scope of publication, and share it with other users.
[1581] A "user" is an individual or organization that uses the system and provides information about events or turning points that have occurred in their life.
[1582] "Input means" refers to the interface and devices that allow users to input information about events and turning points in their lives into the system.
[1583] "Generation means" refers to the function and device by which the server generates video content using an AI algorithm based on input information.
[1584] "Emotion Engine" refers to software and systems for recognizing and analyzing emotions from user responses and provided materials.
[1585] The "display means" refers to a device and interface for visually presenting the generated video content to the user.
[1586] The "publication range setting means" is an interface and device for allowing a user to set the publicity range of video content created by the user.
[1587] The "publication means" refers to a function and device that makes video content public to other users according to a set publicity range.
[1588] "Materials" is a general term for digital data such as photographs, audio data, and music provided by users.
[1589] "Uploading means" refers to the interface and device through which users send materials from their own terminals to the system.
[1590] "Analysis means" refers to a function and device that analyzes uploaded materials and stores them in a database.
[1591] The "image selection means" refers to an interface and device that allows the user to select different video styles (e.g., anime style, drama style, American comic style, etc.) when generating video content.
[1592] "Style change means" refers to functions and devices that automatically change the style of video content based on user selection.
[1593] The present invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, a server uses an AI algorithm to generate video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system is implemented as follows.
[1594] The system works through the website or application that the user accesses. The user first enters their email address and password on the login screen and is authenticated by the server. Once authenticated, the user is redirected to an interface where they can take an AI interview.
[1595] The server checks the user's profile information and uses a generative AI model (e.g., a general-purpose conversation model) to send the user a series of questions. When the user answers questions such as "What was the most memorable event during your time at university?", the server stores the answers in a database and passes them to an emotion engine. The emotion engine uses Google's emotion analysis API or similar to analyze the context and tone of the answers to identify the user's emotional state.
[1596] Next, users upload materials (digital data such as photos, audio files, and music) from their devices through the interface. The devices then send the materials to the server, which analyzes and stores them. Image recognition and audio analysis technologies are used for the analysis.
[1597] Next, the server proposes appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred style from the presented styles. The selected style and options are saved on the server and used for video generation.
[1598] The server combines the selected style, emotional state, and user-provided materials to generate video content using a generative AI model (e.g., DeepMotion's AI animation generation technology). The user can view the generated video on a preview page and request corrections if necessary. The server applies the corrections and regenerates the video.
[1599] Finally, the server displays an interface for setting the visibility of the video content, and the user selects the visibility. The server then makes the video available to other users based on the settings and manages access rights. In this case, the server manages access rights for the video content according to the visibility, allowing for limited or general visibility.
[1600] In addition, the server monitors viewers' reactions to the video content after it has been released, reanalyzing the collected emotional information and saving it as reference information for the next video. This process allows users to create video content that reflects their own emotions and share it with other users within appropriate public limits.
[1601] As a concrete example, consider the case where a user wants to create a video of their college graduation ceremony. First, the user logs in, and the server sends a question such as, "Which event from your graduation ceremony was most memorable?" After the user answers with a specific episode, the server performs sentiment analysis to identify the user's emotions (e.g., joy, excitement, nostalgia). Next, the user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and favorite songs. The server analyzes these materials and recommends video styles based on the sentiment analysis results. The user selects a suggested style, and the server uses AI technology to generate video content based on the selected style and options. The user reviews the generated video and makes any necessary adjustments. Finally, the user selects "private" and shares the video content only with family members. The server manages the video based on the settings, monitors viewer reactions, and uses them to generate the next video. In this way, users can easily create and share emotionally relevant videos of their important events.
[1602] An example of a prompt sentence is asking the user, "Which event from your college graduation ceremony was the most memorable?" In this way, detailed information about the user's experience is collected, and emotion analysis and video generation are performed based on that information.
[1603] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1604] Step 1:
[1605] A user accesses the system's website or app and enters their email address and password on the login screen. The server compares this information with the database, and if authentication is successful, the user is taken to the dashboard screen.
[1606] Input: User's email address and password
[1607] Data processing: Check the input information against the database
[1608] Output: User authentication result, dashboard screen displayed if authentication is successful
[1609] Step 2:
[1610] After logging in, the server checks the user's profile information and uses the generative AI model to initiate an AI interview, sending the user a series of questions to which they respond.
[1611] Input: User profile information, AI model prompt
[1612] Data processing: Generate appropriate questions based on AI algorithms
[1613] Output: Sending interview questions to the user
[1614] Step 3:
[1615] The user answers questions sent by the server. The answers are entered as text or voice. The server receives these answers and stores them in a database.
[1616] Input: User response (text or voice)
[1617] Data processing: Saving response data
[1618] Output: Response data stored in a database
[1619] Step 4:
[1620] The server passes the user's response data to the emotion engine, which uses natural language processing technology to analyze the context and tone of the response and identify the emotional state (e.g., joy, sadness, surprise).
[1621] Input: User response data
[1622] Data processing: Sentiment analysis (using NLP technology)
[1623] Output: Identified emotional state data
[1624] Step 5:
[1625] Users upload digital materials such as photos, audio files, and music from their own devices through the interface. The devices send this data to the server, which then analyzes and stores it in a database.
[1626] Input: User photos, audio files, songs
[1627] Data processing: Analysis and storage of material data
[1628] Output: Material data stored in a database
[1629] Step 6:
[1630] The server displays an interface that suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred video style from the presented options.
[1631] Input: User sentiment analysis results, existing video style data
[1632] Data processing: Proposing appropriate video styles
[1633] Output: User suggestions, user choices
[1634] Step 7:
[1635] The server integrates the selected style, emotional state, and user-provided materials to generate video content using a generative AI model, and the user can view the generated video content on a preview page.
[1636] Input: Selected visual style, emotional state, user material data
[1637] Data processing: Image generation using AI algorithms
[1638] Output: Generated video content
[1639] Step 8:
[1640] The user checks the generated video content and requests corrections if necessary. The server receives the correction request and regenerates the video.
[1641] Input: User's correction request
[1642] Data processing: Regenerating a modified version of the video
[1643] Output: Modified video content
[1644] Step 9:
[1645] The server displays an interface for setting the visibility of video content, and the user selects a setting from options such as private, limited, or public. Based on the setting, the video is made public to other users and access rights are managed.
[1646] Input: User's visibility settings
[1647] Data processing: Applying disclosure settings
[1648] Output: Published video content
[1649] Step 10:
[1650] After the video content is released, the server monitors the viewer's reaction to it, reanalyzes it using the emotion engine, and saves the information as reference for the next video generation.
[1651] Input: Viewer response data (likes, comments, etc.)
[1652] Data processing: sentiment analysis, data storage
[1653] Output: Reference data for future video generation
[1654] In this way, users can create and share video content that reflects their own emotions.
[1655] (Application example 2)
[1656] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1657] In today's world, people have a growing desire to record and share important events and turning points in their lives as video. However, this process requires a lot of time and effort, and it is particularly challenging to generate video that expresses appropriate emotions for each event. Furthermore, determining the extent to which the generated video should be shared and monitoring viewer reactions to collect emotions are also complex. A simple and effective method to solve these problems is needed.
[1658] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input information about events and turning points that occurred in the user's life, means for a generation device to generate video content based on the input information, means for a display device to display the generated video content, means for setting a disclosure range, means for disclosing the video content to other users in accordance with the disclosure range, and means for monitoring viewer reactions and collecting user emotional information. This enables a user to easily generate video content that reflects the emotions of important events, share it within an appropriate disclosure range, and further collect viewer reactions to use in generating the next video.
[1659] "User" refers to an individual or group that uses the system to input information about events and turning points that have occurred in their life, and to generate and share video content.
[1660] "Generation device" refers to a device or software that automatically generates video content based on information input by a user.
[1661] A "display device" is a device or interface for visually displaying generated video content to a user.
[1662] "Means for setting the scope of disclosure" refers to a function or interface that allows users to select and set to whom and to what extent the generated video content will be disclosed.
[1663] "Viewers" refer to other users (third parties) who view published video content.
[1664] "Monitoring" refers to the process of observing and recording audience reactions and behavior, especially capturing audience emotional responses.
[1665] "Emotional information" refers to data or analytical results that indicate the emotional state of a user or viewer.
[1666] "Multiple Materials" refers to digital media files such as photos, audio, and music uploaded by users.
[1667] A "generative AI model" is an artificial intelligence technology that generates a video script based on information input by the user and emotional analysis.
[1668] A "prompt sentence" is an input text given to a generative AI model to instruct it on the content and style of the video content to be generated.
[1669] This invention relates to a system in which users input information about events and turning points in their lives, and AI generates video content based on that information, and an emotion engine is used to recognize and share the user's emotions. This system involves a series of processes, including user information input, preview of the generated video, setting the sharing range, and monitoring viewer reactions.
[1670] System configuration
[1671] User login
[1672] The server provides a means for users to access the system's website or application and enter their login information, which receives the user's authentication information and logs them into the system.
[1673] Conducting AI interviews
[1674] After logging in, the server checks the user's profile information and begins an AI interview. Specifically, it sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[1675] Emotion recognition with emotion engine
[1676] The server then activates an emotion engine based on the user's responses to recognize the user's emotions. The emotion engine analyzes the user's emotional state from their words and tone of voice.
[1677] Provision of materials and analysis
[1678] Users upload materials for video content, such as photos, audio, and music, from their devices. The server stores and analyzes the uploaded materials.
[1679] Selecting the output format
[1680] Users can use an interface to select the style of the video content (e.g., anime, drama, comic, etc.). The server then suggests an appropriate video style based on the results of the emotion engine.
[1681] Image generation and confirmation
[1682] The server uses a generative AI model to generate video content based on user input, provided materials, and recognized emotions. Users can view the generated video on a preview page and request corrections if necessary.
[1683] Setting the visibility and publishing
[1684] The server displays an interface for setting the visibility of the video content, and the user can select the visibility (e.g., private, limited, public). The server makes the video content available to other users based on the settings and manages access rights.
[1685] User response monitoring
[1686] The server monitors the viewer's reaction to the generated video content and collects information on the user's emotions, which is then saved as reference for the next video generation.
[1687] Hardware and software used
[1688] The system is implemented using the following hardware and software:
[1689] Servers: Web servers, database servers
[1690] Device: PC, smartphone, tablet, etc. that users access
[1691] Software: Generative AI models (e.g., OpenAI), emotion analysis engines (e.g., emotion_recognition library)
[1692] Specific examples
[1693] For example, a user may want to film their "college graduation ceremony."
[1694] procedure
[1695] 1. The user logs in and the AI interview begins.
[1696] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[1697] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[1698] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[1699] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[1700] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[1701] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[1702] 8. The user reviews the generated footage and requests corrections if necessary.
[1703] 9. The server completes the edited video.
[1704] 10. The user selects "Private" and sets the video content to be shared only with family members.
[1705] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[1706] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[1707] Prompt Sentence Examples
[1708] For example, the following prompt sentence is input to the generative AI model:
[1709] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[1710] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[1711] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1712] Step 1:
[1713] A user visits the system's website or application and enters their login information. The input to this step is the user's credentials (e.g., email address and password), and the output is a message indicating authentication success or failure. The server checks the user's credentials against information in its database and, if there is a match, allows the user to log in.
[1714] Step 2:
[1715] After logging in, the server checks the user's profile information and starts the AI interview. The input is the user's profile information and a pre-registered question list, and the output is sending questions to the user. The server sequentially sends the user questions such as "Which event at your graduation ceremony was most memorable?"
[1716] Step 3:
[1717] The user answers questions from the server. The input of this step is the user's text answer, and the output is the collected episode details. The user answers specific episodes based on their own experiences, and the information is sent to the server.
[1718] Step 4:
[1719] The server activates an emotion engine based on the collected responses to recognize the user's emotions. The input of this step is the user's response information, and the output is analyzed emotion data. The emotion engine analyzes the language tone and keywords from the user's response to identify the user's emotional state.
[1720] Step 5:
[1721] Users upload materials for video content, such as photos, audio, and music, from their devices to the server. The input is the digital media files (photos, audio, and music) uploaded by the user, and the output is a message confirming the completion of the upload. The device sends the provided materials to the server, which then stores them.
[1722] Step 6:
[1723] The server analyzes the uploaded material and proposes a style for the video content based on it. The input of this step is the uploaded material and emotional data, and the output is a recommended video style. The server analyzes the user's material and emotional state and recommends an appropriate style (e.g., an emotional anime style).
[1724] Step 7:
[1725] The user selects a suggested video style and then selects further options. The input is the style suggestion from the server and the user's selection, and the output is the selected style and options. The user makes their selection through the interface, and that information is sent to the server.
[1726] Step 8:
[1727] The server generates video content using a generative AI model based on the selected style and options. The input for this step is the user's selection information and emotional data, which are converted into specific prompt sentences. The output is the generated video content. The server sends the following prompt sentence to the generative AI model:
[1728] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[1729] Step 9:
[1730] The user can check the generated video on the preview page and request corrections if necessary. The input is the generated video content, and the output is the user's feedback or correction requests. The server displays the generated video to the user and receives any corrections that need to be made.
[1731] Step 10:
[1732] The server completes the revised video. The input of this step is the user's feedback, and the output is the completed video content. The server regenerates the video based on the user's feedback and provides the final version.
[1733] Step 11:
[1734] The user sets the visibility of the completed video content. The input is the visibility selection, and the output is a confirmation message that the settings have been completed. The server centrally makes the video available to other users based on the visibility (private, limited, public) entered.
[1735] Step 12:
[1736] The server monitors viewers' reactions to video content and collects emotional information. The input is viewer reaction data, and the output is analyzed emotional information. The server monitors viewers' reactions and comments and records them as reference information for the next video generation.
[1737] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1738] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1739] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1740] [Fourth embodiment]
[1741] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1742] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1743] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1744] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1745] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1746] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1747] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1748] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1749] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1750] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1751] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1752] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1753] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1754] This invention relates to a system in which users input information about events and turning points that occurred in their lives, and AI generates video content based on that information and shares it with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[1755] System configuration
[1756] 1. User login
[1757] The user is presented with an interface to log into the system, where they enter their login information and are authenticated.
[1758] 2. Conducting AI interviews
[1759] After logging in, the server checks the user's profile information and begins the AI interview.
[1760] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[1761] 3. Provision of Materials
[1762] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[1763] The server stores and analyzes the uploaded material.
[1764] 4. Select output format
[1765] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[1766] Users can also select additional options such as trends and historical background of the time.
[1767] 5. Image generation and confirmation
[1768] The server uses AI algorithms to generate video content based on user input and provided materials.
[1769] Users can check the generated video on the preview page and request corrections if necessary.
[1770] 6. Setting the visibility and publishing
[1771] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[1772] The server makes the video content available to other users and manages access rights based on the settings.
[1773] Specific examples
[1774] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[1775] 1. The user logs in and the AI interview begins.
[1776] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[1777] 3. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[1778] 4. The server analyzes the user's material and recommends a video style (for example, anime style) based on that analysis.
[1779] 5. The user selects an anime style and then selects the "Trend of the Year" option.
[1780] 6. The server generates the video content based on the selected style and options.
[1781] 7. The user reviews the generated footage and requests corrections if necessary.
[1782] 8. The server completes the edited video.
[1783] 9. The user selects "Private" and sets the video content to be shared only with family members.
[1784] 10. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[1785] The system allows users to record important events in their lives and easily share them with others while respecting privacy.
[1786] The processing flow will be explained below.
[1787] Step 1:
[1788] A user visits the system's website or app and enters their login information.
[1789] Step 2:
[1790] The server authenticates the login information and verifies the user's profile information.
[1791] Step 3:
[1792] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[1793] Step 4:
[1794] The user enters text responses to the AI interview questions.
[1795] Step 5:
[1796] The server analyzes the user's answers and generates and displays the next question (e.g., "What is the duration of the event?").
[1797] Step 6:
[1798] The user answers additional questions as the interview continues.
[1799] Step 7:
[1800] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[1801] Step 8:
[1802] The server receives the uploaded material and stores it in storage.
[1803] Step 9:
[1804] The server analyzes the stored material and extracts the information necessary to generate the video.
[1805] Step 10:
[1806] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[1807] Step 11:
[1808] The user selects their preferred style and options (such as what was popular at the time).
[1809] Step 12:
[1810] The server stores the user's selections and prepares the video for generation.
[1811] Step 13:
[1812] The server uses AI algorithms to generate video content based on user input and provided materials.
[1813] Step 14:
[1814] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[1815] Step 15:
[1816] The user can check the footage on the preview page and enter correction requests if necessary.
[1817] Step 16:
[1818] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[1819] Step 17:
[1820] The user makes a final confirmation and sets the public range (private, limited, public).
[1821] Step 18:
[1822] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[1823] Example 1
[1824] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1825] In modern society, there are limited ways to record special events and important moments in video and share them with others. As a result, users face challenges in easily creating and sharing video content, requiring significant effort and technical knowledge. Furthermore, there is a lack of flexible ways to change the style and image of video content, making it difficult to meet the diverse needs of users. Furthermore, from the perspective of privacy protection, a lack of systems that allow users to easily set and manage the scope of disclosure has been pointed out.
[1826] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1827] In this invention, the server includes means for a user to input information about events or important events in his or her life, means for generating and presenting questions to the user and collecting the information input by the user, means for uploading multiple digital data provided by the user, means for analyzing and saving the uploaded digital data, means for a generation device to generate video content based on the collected information and analyzed data, means for a display device to display the generated video content, means for setting a disclosure range, and means for making the video content available to other users in accordance with the disclosure range. This enables users, without special knowledge, to easily visualize their important events, express them in various styles, and share them with others while protecting their privacy.
[1828] A "user" is someone who uses the system to visualize their own events or important events and share them with others.
[1829] "Server" refers to the central computer that controls and manages the entire system, and analyzes and stores information entered by users and uploaded material data.
[1830] "Terminal" means a device that allows a user to access the system, input information, and view video content.
[1831] A "generation device" is a device or software that uses AI algorithms to generate video content based on information collected from users and materials provided by them.
[1832] A "display device" is a device for displaying generated video content to a user.
[1833] The "means for generating questions" is a function for generating appropriate questions based on information input by the user and presenting them to the user.
[1834] "Digital data" refers to multimedia materials such as photographs, audio, and music provided by users.
[1835] "Means for uploading" refers to the functionality that allows users to transfer digital data to the system.
[1836] "Means for analyzing and storing" refers to the function for analyzing uploaded digital data and storing it within the system.
[1837] The "means for setting the disclosure range" is a function for selecting and setting the disclosure range of the generated video content.
[1838] The "means for publishing according to the public range" is a function for making video content accessible to other users based on the public range that has been set.
[1839] This invention relates to a system in which users input information about events or important events in their lives, and AI generates video content based on that information and shares it with other users. This system mainly involves information exchange and processing between a server, a terminal, and the user.
[1840] System configuration
[1841] 1. User login
[1842] To access the system, the terminal displays a login screen. The user enters their email address and password, which the terminal sends to the server. The server authenticates them by looking up their email address in a database (e.g., MongoDB or MySQL), and if authentication is successful, allows the user to proceed to the next step.
[1843] 2. Conducting AI interviews
[1844] The server retrieves the profile information of the user who successfully logged in and generates appropriate questions using a natural language processing model (e.g., GPT-4). The server sends a series of questions to the device, which displays them to the user. The user answers the questions, and the device sends the answers to the server.
[1845] 3. Provision of Materials
[1846] Users use their devices to upload digital data such as their photos, audio files, and songs. The devices send this data to a server, which then analyzes the data using image processing libraries (e.g., OpenCV) and audio analysis libraries (e.g., Librosa) and stores it in a storage system (e.g., Amazon S3).
[1847] 4. Select output format
[1848] The device displays an interface for selecting the style and image of the video content, allowing the user to choose styles such as anime, TV drama, American comics, etc. Additional options such as the trends of the time and historical background can also be selected. The device sends the selection results to the server, which stores them in a database.
[1849] 5. Image generation and confirmation
[1850] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials. The server temporarily saves the generated video and sends a URL to the device. The device displays a video preview screen, which the user can check. If the user sends a correction request as needed, the server accepts the correction request and regenerates the video.
[1851] 6. Setting the visibility and publishing
[1852] The device displays an interface for setting the visibility of the video content. The user can select from options such as private, limited, or public. The device sends the selection result to the server, which then manages access rights for the video content based on the selected visibility.
[1853] Specific examples
[1854] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[1855] 1. The user logs in and the system starts the AI interview.
[1856] 2. The server sends the user a question such as, "Which event from your graduation ceremony was most memorable?", and the user answers with a specific episode.
[1857] 3. The user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and their favorite songs from their device.
[1858] 4. The server analyzes the uploaded material and suggests a video style (e.g., anime style).
[1859] 5. The user selects an anime style and chooses "Trend of the Year" as an additional option.
[1860] 6. The server generates the video content based on the selected style and options, using DeepArt or StyleGAN.
[1861] 7. The user checks the generated video on their device and requests corrections if necessary.
[1862] 8. The server completes the edited video.
[1863] 9. The user selects "Private" and sets the video content to be shared only with family members.
[1864] 10. The server publishes the video based on the user's settings and manages it so that only users within the specified range of access can access the video.
[1865] Example prompts to input to the generative AI model
[1866] A possible prompt for a generative AI model (e.g., GPT-4) might look something like this:
[1867] The user is talking about their college graduation ceremony. Use the following elements to generate animated video content:
[1868] Photo 1: A photo of me with my friends at the graduation ceremony.
[1869] Photo 2: Exterior of the ceremony venue.
[1870] Audio file: Audio of the speech on graduation day.
[1871] Song: A memorable song from graduation ceremony.
[1872] The tone of the video should be emotive and reflect the trends of the time.
[1873] This system allows users, without any special knowledge, to easily visualize their important events, express them in a variety of styles, and share them with others while protecting their privacy.
[1874] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1875] System program processing flow
[1876] Step 1:
[1877] User login
[1878] The device displays the login screen.
[1879] Input: User's email address and password.
[1880] Specific operation: Display a login form using HTML / CSS / JavaScript.
[1881] The user enters their email address and password.
[1882] Output: The login information entered.
[1883] The terminal sends the input information to the server.
[1884] Input: The login information entered.
[1885] Output: HTTP request to the server.
[1886] The server performs authentication by referencing a database (e.g. MongoDB or MySQL).
[1887] Input: The login information sent to the server.
[1888] Specific behavior: Provides an API (e.g., RESTful API) on the backend and handles authentication logic.
[1889] Output: Authentication result (success / failure).
[1890] The server returns the authentication result to the terminal, and if successful, proceeds to the next step.
[1891] Input: Authentication result.
[1892] Output: Next screen based on authentication result (instructions to proceed to next step if successful).
[1893] Step 2:
[1894] Conducting AI interviews
[1895] The server retrieves the user's profile information after a successful login.
[1896] Input: Login success notification.
[1897] Specific Actions: Retrieve user profile from database.
[1898] Output: Profile information.
[1899] The server uses the profile information to generate appropriate questions using a natural language processing model (e.g., GPT-4).
[1900] Input: Profile information.
[1901] What it does: Generate questions using GPT-4.
[1902] Output: Question list.
[1903] The server sends the generated question to the terminal.
[1904] Input: Questionnaire.
[1905] Output: HTTP response containing the questionnaire.
[1906] The device displays a question and the user enters the answer.
[1907] Input: Questionnaire.
[1908] Specific operation: Dynamically generate and display a question form.
[1909] Output: The user's answer.
[1910] The device sends the user's answer to the server.
[1911] Input: The user's answer.
[1912] Output: HTTP request to the server (response data).
[1913] Step 3:
[1914] Provision of materials
[1915] The device will display the interface for uploading materials.
[1916] Input: The upload request.
[1917] Specific behavior: HTML <input type="file"> Implement file selection using an element.
[1918] Output: Upload screen.
[1919] Users select and upload digital data such as photos, audio, and music.
[1920] Input: Materials (photos, audio, music, etc.).
[1921] Output: The digital data to be uploaded.
[1922] The device sends the uploaded data to the server.
[1923] Input: Digital data to be uploaded.
[1924] Output: HTTP request to server (digital data).
[1925] The server receives the data and stores it in a repository.
[1926] Input: Uploaded digital data.
[1927] Specific behavior: Stores data in a file storage system (e.g., Amazon S3).
[1928] Output: Stored digital data.
[1929] The server analyzes the stored data.
[1930] Input: Stored digital data.
[1931] Specific operation: Analyze the data using an image processing library (e.g., OpenCV) or a sound analysis library (e.g., Librosa).
[1932] Output: Analysis results.
[1933] Step 4:
[1934] Selecting the output format
[1935] The terminal displays an interface for selecting the style of the video content.
[1936] Input: A style selection request.
[1937] What it does: Provide users with choices using drop-down menus or radio buttons.
[1938] Output: Style selection screen.
[1939] The user selects a style and image from the options provided.
[1940] Enter: Style selection.
[1941] Output: Selected style information.
[1942] The terminal transmits the selection result to the server.
[1943] Input: Selected style information.
[1944] Output: HTTP request to the server (style information).
[1945] The server stores the user's selection in a database.
[1946] Input: Selected style information.
[1947] Specific action: Save to database.
[1948] Output: The saved style information.
[1949] Step 5:
[1950] Image generation and confirmation
[1951] The server generates video content using an AI algorithm (e.g., DeepArt or StyleGAN) based on the user's input information and provided materials.
[1952] Input: User input information and materials provided.
[1953] Specific operation: Generate images using DeepArt and StyleGAN.
[1954] Output: The generated video content.
[1955] The server temporarily stores the generated video and sends its URL to the terminal.
[1956] Input: Generated video content.
[1957] Output: URL of the archived video.
[1958] The device displays a video preview screen, and the user checks the video.
[1959] Input: The URL of the downloaded video.
[1960] Specific behavior: HTML <video>Display video using tags.
[1961] Output: The confirmation status of the user.
[1962] The user sends a modification request to the server as needed.
[1963] Input: The user's correction request.
[1964] Output: HTTP request to the server (modification request).
[1965] Step 6:
[1966] Setting the visibility and publishing
[1967] The device displays an interface for selecting the disclosure range.
[1968] Input: Disclosure scope setting request.
[1969] Specific behavior: Provide options for the scope of disclosure using radio buttons or checkboxes.
[1970] Output: Public range selection screen.
[1971] The user chooses the scope of disclosure.
[1972] Input: Select the disclosure range.
[1973] Output: Selected disclosure information.
[1974] The terminal transmits the selection result to the server.
[1975] Input: Selected disclosure information.
[1976] Output: HTTP request to the server (public range information).
[1977] The server sets the video content to be made public based on the selected public range.
[1978] Input: Selected disclosure information.
[1979] Specific operation: Set the scope of disclosure using ACL (Access Control List).
[1980] Output: The set disclosure range.
[1981] The server manages access rights for other users based on the public settings.
[1982] Input: The set disclosure range.
[1983] Output: Managed access rights.
[1984] (Application example 1)
[1985] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1986] In today's virtual stores, it is difficult to provide a personalized shopping experience based on users' emotions and personal memories. Therefore, a system that allows users to enjoy shopping while feeling an emotional connection is needed. Furthermore, existing systems lack the means to effectively utilize user-provided materials, display them as video content, and provide product suggestions and special offers. Furthermore, the technology to visualize individual user events and turning points and provide them as moving experiences within the virtual store is underdeveloped.
[1987] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1988] In this invention, the server includes: means for a user to input information about events and turning points that occurred in the user's life; means for a generation device to generate video content based on the input information; means for a display device to display the generated video content; means for setting a disclosure range; means for disclosing the video content to other users according to the disclosure range; and means for displaying products and special offers using the generated video content so that the user can personalize an emotional shopping experience in the virtual store. This allows the user to enjoy a virtual shopping experience based on personal memories while feeling an emotional connection, and further allows the user to effectively receive product suggestions and special offers through the video content.
[1989] "User" refers to an individual who uses the System.
[1990] "Events" refer to memorable experiences or turning points that occurred in the user's life.
[1991] A "turning point" refers to a time or event in a user's life when an important change or decision was made.
[1992] "Means for inputting information" refers to the interface through which a user provides information about their event to the system.
[1993] "Generation device" refers to a computer program and hardware system that creates video content based on input information.
[1994] "Display device" refers to a device that visually presents generated video content to a user.
[1995] The "means for setting the public range" refers to an interface that allows a user to select the viewing range of the generated video content.
[1996] "Publicity" refers to the range of users who can view the generated video content.
[1997] "Virtual store" refers to a virtual shopping environment provided on the Internet.
[1998] "Shopping Experience" refers to the series of activities in which a user browses, selects, and purchases products within a virtual store.
[1999] "Means for displaying products and special offers" refers to an interface for presenting related products and discount information within the generated video content.
[2000] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content and displays products and special offers in a virtual store. Specific embodiments of the system are described below.
[2001] System configuration
[2002] 1. User login
[2003] Users use the interface to log in to the system. When the user enters their login information, the server performs authentication.
[2004] 2. Conducting AI interviews
[2005] After logging in, the server checks the user's profile information and begins an AI interview. The server sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[2006] 3. Provision of Materials
[2007] Users upload digital data such as photos, audio, and favorite music through their devices, and the server stores and analyzes the uploaded material.
[2008] 4. Select output format
[2009] The server provides an interface for users to select the style and image of the video content (e.g., anime, TV drama, American comic book, etc.) Users can also select additional options such as the trends of the time and the historical background.
[2010] 5. Video Generation and Application to Virtual Stores
[2011] The server uses a generative AI model to generate video content based on user input and provided materials, which is then used to display products and special offers within the virtual store.
[2012] Hardware and software used
[2013] Hardware: smartphones, tablets, servers
[2014] Software: Online platform APIs, generative AI models (e.g., OpenAI GPT-3)
[2015] Data processing: Analyzes user input data (text, images, audio) and performs processes ranging from prompt generation to video content generation.
[2016] Specific examples
[2017] For example, consider the case where a user wants to create a video of their "college graduation ceremony." After logging in, the user answers questions such as "Which event from your graduation ceremony was most memorable?" through an AI interview. The server analyzes the materials provided by the user, such as photos, audio files, and favorite songs, and recommends video styles based on these. If the user selects an anime-style style and also selects "Trend of the Year" as an option, the server uses a generative AI model to generate video content based on the selected style and options. This video is then used to display products and special offers within the virtual store.
[2018] Prompt Sentence Examples
[2019] text
[2020] Title: College Graduation
[2021] Details: Wonderful days spent with many friends, and a graduation day that will remain in my memory forever.
[2022] Image files: ["graduation_photo1.jpg", "graduation_photo2.jpg"]
[2023] Audio file: ["graduation_speech.mp3"]
[2024] Style: Emotional documentary
[2025] Generate effective footage.
[2026] In this way, a system is realized that provides an emotional shopping experience in a virtual store based on the user's special events.
[2027] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2028] Step 1:
[2029] A user logs in to the system.
[2030] Input: User login information (username, password).
[2031] Output: If authentication is successful, proceed to next step. If not, prompt to retry.
[2032] Specific operation: The user enters the username and password into the login interface, and the server authenticates it. If the authentication is successful, the server obtains the user's profile information and proceeds to the next step.
[2033] Step 2:
[2034] AI interview begins.
[2035] Input: User profile information and login state.
[2036] Output: Detailed information about the user's events and turning points.
[2037] What it does: After logging in, the server asks the user a series of questions, which the user answers, and the server collects these answers, recording details of events and turning points that the user enters during this process.
[2038] Step 3:
[2039] Users upload materials.
[2040] Input: Digital data provided by the user, such as photos, audio, or favorite songs.
[2041] Output: Uploaded material is stored on the server and analyzed.
[2042] How it works: Users use their devices to upload photos, audio files, music, and other content to a server. The server then analyzes and stores the data. Image recognition and audio analysis algorithms are used for the analysis.
[2043] Step 4:
[2044] Video content style selection.
[2045] Input: User's event details and uploaded materials.
[2046] Output: User-selected video style and additional options.
[2047] Specific operation: Through the interface, the server allows the user to select the style of the video content (e.g., anime, drama, American comics, etc.). Additional options include the selection of the trends of the time and the historical background.
[2048] Step 5:
[2049] Video content generation.
[2050] Input: User event details, uploaded footage, selected video style and additional options.
[2051] Output: The generated video content.
[2052] Specific operation: The server uses the generative AI model to generate video content based on the user's input information and materials, inputting prompt sentences into the generative AI model and outputting video data based on them.
[2053] Step 6:
[2054] Displaying products and special offers in a virtual store.
[2055] Input: Generated video content.
[2056] Output: Products and special offers displayed in a virtual store with emotional video content.
[2057] What it does: The generated video content is played in the virtual store, where product suggestions and special offers related to the user's current situation are displayed, allowing the user to browse and purchase products with an emotional connection.
[2058] Step 7:
[2059] Setting the visibility of video content and publishing it.
[2060] Input: Generated video content, user-selected visibility.
[2061] Output: Video content with access rights controlled for other users based on the set visibility.
[2062] Specific operation: The server provides an interface for users to set the visibility of video content (private, limited, public, etc.). Based on the visibility set by the user, the server makes the video available to other users and manages access rights.
[2063] Through these steps, the system can visualize the user's special events and provide an emotional shopping experience within the virtual store.
[2064] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2065] This invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, AI generates video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. This system exchanges and processes information between a server, terminals, and users, and is implemented as follows:
[2066] System configuration
[2067] 1. User login
[2068] A user visits the system's website or app and enters their login information.
[2069] 2. Conducting AI interviews
[2070] After logging in, the server checks the user's profile information and begins the AI interview.
[2071] The server sends the user a series of questions, which the user answers to gather details about the events or turning points they want to visualize.
[2072] 3. Emotion recognition using the emotion engine
[2073] The server activates an emotion engine to recognize the user's emotion from the content of the user's answers and the provided materials.
[2074] The emotion engine analyzes the user's emotional state based on the context, words used, tone of voice, etc.
[2075] 4. Provision of Materials
[2076] The device allows users to upload digital data such as photos, audio, and favorite music provided by the user.
[2077] The server stores and analyzes the uploaded material.
[2078] 5. Select output format
[2079] An interface is displayed that allows the user to select the style or image of the video content (e.g., anime style, drama style, American comic style, etc.).
[2080] The server proposes an appropriate video style based on the results of the emotion engine.
[2081] 6. Image generation and confirmation
[2082] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[2083] Users can check the generated video on the preview page and request corrections if necessary.
[2084] 7. Setting the visibility and publishing
[2085] The server displays an interface for users to set the visibility of their video content, with options such as private, limited, or public.
[2086] The server makes the video content available to other users and manages access rights based on the settings.
[2087] 8. User response monitoring
[2088] The server monitors the user's reaction to the generated video content and collects emotional information about the user.
[2089] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[2090] Specific examples
[2091] For example, let us consider a case where a user wants to visualize his / her "graduation ceremony from college."
[2092] 1. The user logs in and the AI interview begins.
[2093] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[2094] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[2095] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[2096] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[2097] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[2098] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[2099] 8. The user reviews the generated footage and requests corrections if necessary.
[2100] 9. The server completes the edited video.
[2101] 10. The user selects "Private" and sets the video content to be shared only with family members.
[2102] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[2103] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[2104] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[2105] The processing flow will be explained below.
[2106] Step 1:
[2107] A user visits the system's website or app and enters their login information.
[2108] Step 2:
[2109] The server authenticates the login information and verifies the user's profile information.
[2110] Step 3:
[2111] The server displays the initial screen of the AI interview and asks the user questions about the event they want to visualize (e.g., "What event do you want to visualize?").
[2112] Step 4:
[2113] The user enters text responses to the AI interview questions.
[2114] Step 5:
[2115] The server receives the user's response and activates the emotion engine to analyze the user's emotional state. During this process, emotions are recognized from the content and context of the user's response, the words used, and the tone of voice.
[2116] Step 6:
[2117] The server stores the analysis results of the emotion engine and generates and displays the next question (e.g., "What is the duration of the event?").
[2118] Step 7:
[2119] The user answers additional questions as the interview continues.
[2120] Step 8:
[2121] Users upload the materials needed to create a video (photos, audio, favorite music, etc.).
[2122] Step 9:
[2123] The server receives the uploaded material and stores it in storage.
[2124] Step 10:
[2125] The server analyzes the stored material and extracts the information necessary to generate the video.
[2126] Step 11:
[2127] The server displays a screen for selecting the output format, allowing the user to select the style or image of the video (e.g., anime, drama, American comics, etc.).
[2128] Step 12:
[2129] The user selects their preferred style and options (such as what was popular at the time).
[2130] Step 13:
[2131] The server presents a suggested output format (e.g., an emotional anime style) based on the analysis results of the emotion engine.
[2132] Step 14:
[2133] The server stores the user's selections and prepares the video for generation.
[2134] Step 15:
[2135] The server uses AI algorithms to generate video content based on user input, provided materials, and recognized emotions.
[2136] Step 16:
[2137] The server displays the generated video content on a preview page and sends a confirmation request notification to the user.
[2138] Step 17:
[2139] The user can check the footage on the preview page and enter correction requests if necessary.
[2140] Step 18:
[2141] The server receives the user's correction request, makes the necessary corrections, and regenerates the video.
[2142] Step 19:
[2143] The user makes a final confirmation and sets the public range (private, limited, public).
[2144] Step 20:
[2145] The server publishes video content based on the user's settings and performs access management for users within the specified range of disclosure.
[2146] Step 21:
[2147] The server monitors viewers' reactions to the published video content and collects user emotional information.
[2148] Step 22:
[2149] The emotion engine analyzes the collected emotional information and saves it as reference information for the next video generation.
[2150] Example 2
[2151] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2152] Today's users have a growing need to digitize the events and turning points in their lives and save and share them as emotionally charged video content. However, existing systems for achieving this lacked the technology to properly analyze users' emotions and reflect them in the video content. Furthermore, they lacked a means to generate video content with different images and allow users to easily set the visibility of the content, which hindered the user experience.
[2153] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input information about events and turning points that occurred in the user's life, a means for the server to generate video content using an AI algorithm based on the input information, and a means for an emotion engine to recognize emotions from the user's responses and provided materials. This enables a user to generate video content that reflects their own emotions, select a video style with a different image, easily set the scope of publication, and share it with other users.
[2154] A "user" is an individual or organization that uses the system and provides information about events or turning points that have occurred in their life.
[2155] "Input means" refers to the interface and devices that allow users to input information about events and turning points in their lives into the system.
[2156] "Generation means" refers to the function and device by which the server generates video content using an AI algorithm based on input information.
[2157] "Emotion Engine" refers to software and systems for recognizing and analyzing emotions from user responses and provided materials.
[2158] The "display means" refers to a device and interface for visually presenting the generated video content to the user.
[2159] The "publication range setting means" is an interface and device for allowing a user to set the publicity range of video content created by the user.
[2160] The "publication means" refers to a function and device that makes video content public to other users according to a set publicity range.
[2161] "Materials" is a general term for digital data such as photographs, audio data, and music provided by users.
[2162] "Uploading means" refers to the interface and device through which users send materials from their own terminals to the system.
[2163] "Analysis means" refers to a function and device that analyzes uploaded materials and stores them in a database.
[2164] The "image selection means" refers to an interface and device that allows the user to select different video styles (e.g., anime style, drama style, American comic style, etc.) when generating video content.
[2165] "Style change means" refers to functions and devices that automatically change the style of video content based on user selection.
[2166] The present invention relates to a system in which a user inputs information about events and turning points that occurred in their life, and based on that information, a server uses an AI algorithm to generate video content, recognizes the user's emotions using an emotion engine, and shares the content with other users. The system is implemented as follows.
[2167] The system works through the website or application that the user accesses. The user first enters their email address and password on the login screen and is authenticated by the server. Once authenticated, the user is redirected to an interface where they can take an AI interview.
[2168] The server checks the user's profile information and uses a generative AI model (e.g., a general-purpose conversation model) to send the user a series of questions. When the user answers questions such as "What was the most memorable event during your time at university?", the server stores the answers in a database and passes them to an emotion engine. The emotion engine uses Google's emotion analysis API or similar to analyze the context and tone of the answers to identify the user's emotional state.
[2169] Next, users upload materials (digital data such as photos, audio files, and music) from their devices through the interface. The devices then send the materials to the server, which analyzes and stores them. Image recognition and audio analysis technologies are used for the analysis.
[2170] Next, the server suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred style from the presented styles. The selected style and options are saved on the server and used for video generation.
[2171] The server combines the selected style, emotional state, and user-provided materials to generate video content using a generative AI model (e.g., DeepMotion's AI animation generation technology). The user can view the generated video on a preview page and request corrections if necessary. The server applies the corrections and regenerates the video.
[2172] Finally, the server displays an interface for setting the visibility of the video content, and the user selects the visibility. The server then makes the video available to other users based on the settings and manages access rights. In this case, the server manages access rights for the video content according to the visibility, allowing for limited or general visibility.
[2173] In addition, the server monitors viewers' reactions to the video content after it has been released, reanalyzing the collected emotional information and saving it as reference information for the next video. This process allows users to create video content that reflects their own emotions and share it with other users within appropriate public limits.
[2174] As a concrete example, consider the case where a user wants to create a video of their college graduation ceremony. First, the user logs in, and the server sends a question such as, "Which event from your graduation ceremony was most memorable?" After the user answers with a specific episode, the server performs sentiment analysis to identify the user's emotions (e.g., joy, excitement, nostalgia). Next, the user uploads photos taken at the graduation ceremony, audio files of conversations with friends, and favorite songs. The server analyzes these materials and recommends video styles based on the sentiment analysis results. The user selects a suggested style, and the server uses AI technology to generate video content based on the selected style and options. The user reviews the generated video and makes any necessary adjustments. Finally, the user selects "private" and shares the video content only with family members. The server manages the video based on the settings, monitors viewer reactions, and uses them to generate the next video. In this way, users can easily create and share emotionally relevant videos of their important events.
[2175] An example of a prompt sentence is asking the user, "Which event from your college graduation ceremony was the most memorable?" In this way, detailed information about the user's experience is collected, and emotion analysis and video generation are performed based on that information.
[2176] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2177] Step 1:
[2178] A user accesses the system's website or app and enters their email address and password on the login screen. The server compares this information with the database, and if authentication is successful, the user is taken to the dashboard screen.
[2179] Input: User's email address and password
[2180] Data processing: Check the input information against the database
[2181] Output: User authentication result, dashboard screen displayed if authentication is successful
[2182] Step 2:
[2183] After logging in, the server checks the user's profile information and uses the generative AI model to initiate an AI interview, sending the user a series of questions to which they respond.
[2184] Input: User profile information, AI model prompt
[2185] Data processing: Generate appropriate questions based on AI algorithms
[2186] Output: Sending interview questions to the user
[2187] Step 3:
[2188] The user answers questions sent by the server. The answers are entered as text or voice. The server receives these answers and stores them in a database.
[2189] Input: User response (text or voice)
[2190] Data processing: Saving response data
[2191] Output: Response data stored in a database
[2192] Step 4:
[2193] The server passes the user's response data to the emotion engine, which uses natural language processing technology to analyze the context and tone of the response and identify the emotional state (e.g., joy, sadness, surprise).
[2194] Input: User response data
[2195] Data processing: Sentiment analysis (using NLP technology)
[2196] Output: Identified emotional state data
[2197] Step 5:
[2198] Users upload digital materials such as photos, audio files, and music from their own devices through the interface. The devices send this data to the server, which then analyzes and stores it in a database.
[2199] Input: User photos, audio files, songs
[2200] Data processing: Analysis and storage of material data
[2201] Output: Material data stored in a database
[2202] Step 6:
[2203] The server displays an interface that suggests appropriate video styles (e.g., anime style, drama style) to the user based on the results of the emotion engine. The user selects their preferred video style from the presented options.
[2204] Input: User sentiment analysis results, existing video style data
[2205] Data processing: Proposing appropriate video styles
[2206] Output: User suggestions, user choices
[2207] Step 7:
[2208] The server integrates the selected style, emotional state, and user-provided materials to generate video content using a generative AI model, and the user can view the generated video content on a preview page.
[2209] Input: Selected visual style, emotional state, user material data
[2210] Data processing: Image generation using AI algorithms
[2211] Output: Generated video content
[2212] Step 8:
[2213] The user checks the generated video content and requests corrections if necessary. The server receives the correction request and regenerates the video.
[2214] Input: User's correction request
[2215] Data processing: Regenerating a modified version of the video
[2216] Output: Modified video content
[2217] Step 9:
[2218] The server displays an interface for setting the visibility of video content, and the user selects a setting from options such as private, limited, or public. Based on the setting, the video is made public to other users and access rights are managed.
[2219] Input: User's visibility settings
[2220] Data processing: Applying disclosure settings
[2221] Output: Published video content
[2222] Step 10:
[2223] After the video content is released, the server monitors the viewer's reaction to it, reanalyzes it using the emotion engine, and saves the information as reference for the next video generation.
[2224] Input: Viewer response data (likes, comments, etc.)
[2225] Data processing: sentiment analysis, data storage
[2226] Output: Reference data for future video generation
[2227] In this way, users can create and share video content that reflects their own emotions.
[2228] (Application example 2)
[2229] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2230] In today's world, people have a growing desire to record and share important events and turning points in their lives as video. However, this process requires a lot of time and effort, and it is particularly challenging to generate video that expresses appropriate emotions for each event. Furthermore, determining the extent to which the generated video should be shared and monitoring viewer reactions to collect emotions are also complex. A simple and effective method to solve these problems is needed.
[2231] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input information about events and turning points that occurred in the user's life, means for a generation device to generate video content based on the input information, means for a display device to display the generated video content, means for setting a disclosure range, means for disclosing the video content to other users in accordance with the disclosure range, and means for monitoring viewer reactions and collecting user emotional information. This enables a user to easily generate video content that reflects the emotions of important events, share it within an appropriate disclosure range, and further collect viewer reactions to use in generating the next video.
[2232] "User" refers to an individual or group that uses the system to input information about events and turning points that have occurred in their life, and to generate and share video content.
[2233] "Generation device" refers to a device or software that automatically generates video content based on information input by a user.
[2234] A "display device" is a device or interface for visually displaying generated video content to a user.
[2235] "Means for setting the scope of disclosure" refers to a function or interface that allows users to select and set to whom and to what extent the generated video content will be disclosed.
[2236] "Viewers" refer to other users (third parties) who view published video content.
[2237] "Monitoring" refers to the process of observing and recording audience reactions and behavior, especially capturing audience emotional responses.
[2238] "Emotional information" refers to data or analytical results that indicate the emotional state of a user or viewer.
[2239] "Multiple Materials" refers to digital media files such as photos, audio, and music uploaded by users.
[2240] A "generative AI model" is an artificial intelligence technology that generates a video script based on information input by the user and emotional analysis.
[2241] A "prompt sentence" is an input text given to a generative AI model to instruct it on the content and style of the video content to be generated.
[2242] This invention relates to a system in which users input information about events and turning points that occurred in their lives, and AI generates video content based on that information, and an emotion engine is used to recognize and share the user's emotions. This system involves a series of processes, including user information input, preview of the generated video, setting the sharing range, and monitoring viewer reactions.
[2243] System configuration
[2244] User login
[2245] The server provides a means for users to access the system's website or application and enter their login information, which receives the user's authentication information and logs them into the system.
[2246] Conducting AI interviews
[2247] After logging in, the server checks the user's profile information and begins an AI interview. Specifically, it sends the user a series of questions, and as the user answers them, it collects detailed information about the events and turning points they want to visualize.
[2248] Emotion recognition with emotion engine
[2249] The server then activates an emotion engine based on the user's responses to recognize the user's emotions. The emotion engine analyzes the user's emotional state from their words and tone of voice.
[2250] Provision of materials and analysis
[2251] Users upload materials for video content, such as photos, audio, and music, from their devices. The server stores and analyzes the uploaded materials.
[2252] Selecting the output format
[2253] Users can use an interface to select the style of the video content (e.g., anime, drama, comic, etc.). The server then suggests an appropriate video style based on the results of the emotion engine.
[2254] Image generation and confirmation
[2255] The server uses a generative AI model to generate video content based on user input, provided materials, and recognized emotions. Users can view the generated video on a preview page and request corrections if necessary.
[2256] Setting the visibility and publishing
[2257] The server displays an interface for setting the visibility of the video content, and the user can select the visibility (e.g., private, limited, public). The server makes the video content available to other users based on the settings and manages access rights.
[2258] User response monitoring
[2259] The server monitors the viewer's reaction to the generated video content and collects information on the user's emotions, which is then saved as reference for the next video generation.
[2260] Hardware and software used
[2261] The system is implemented using the following hardware and software:
[2262] Servers: Web servers, database servers
[2263] Device: PC, smartphone, tablet, etc. that users access
[2264] Software: Generative AI models (e.g., OpenAI), emotion analysis engines (e.g., emotion_recognition library)
[2265] Specific examples
[2266] For example, a user may want to film their "college graduation ceremony."
[2267] procedure
[2268] 1. The user logs in and the AI interview begins.
[2269] 2. The server sends a question such as "Which event at your graduation ceremony was most memorable?" and the user answers with a specific episode.
[2270] 3. The server activates an emotion engine based on the user's answers and analyzes the emotions experienced by the user (e.g., joy, excitement, nostalgia).
[2271] 4. Users upload photos taken at the graduation ceremony, audio files of conversations with friends, their favorite songs, etc.
[2272] 5. The server analyzes the user's material and recommends a video style (e.g., an emotional anime style) based on that analysis.
[2273] 6. The user selects a suggested style and then selects the "Trend of the Year" option.
[2274] 7. The server generates video content based on the selected style and options, integrating the analysis results of the emotion engine.
[2275] 8. The user reviews the generated footage and requests corrections if necessary.
[2276] 9. The server completes the edited video.
[2277] 10. The user selects "Private" and sets the video content to be shared only with family members.
[2278] 11. The server makes the video public based on the user's settings and manages the video so that only users within the specified public range can access it.
[2279] 12. The server monitors the viewer's reaction to the video content and uses the collected emotional information as reference information when generating the next video.
[2280] Prompt Sentence Examples
[2281] For example, the following prompt sentence is input to the generative AI model:
[2282] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[2283] This system allows users to visualize their important events, easily generate content that reflects their emotions, and share it within appropriate public boundaries.
[2284] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2285] Step 1:
[2286] A user visits the system's website or application and enters their login information. The input to this step is the user's credentials (e.g., email address and password), and the output is a message indicating authentication success or failure. The server checks the user's credentials against information in its database and, if there is a match, allows the user to log in.
[2287] Step 2:
[2288] After logging in, the server checks the user's profile information and starts the AI interview. The input is the user's profile information and a pre-registered question list, and the output is sending questions to the user. The server sequentially sends the user questions such as "Which event at your graduation ceremony was most memorable?"
[2289] Step 3:
[2290] The user answers questions from the server. The input of this step is the user's text answer, and the output is the collected episode details. The user answers specific episodes based on their own experiences, and the information is sent to the server.
[2291] Step 4:
[2292] The server activates an emotion engine based on the collected responses to recognize the user's emotions. The input of this step is the user's response information, and the output is analyzed emotion data. The emotion engine analyzes the language tone and keywords from the user's response to identify the user's emotional state.
[2293] Step 5:
[2294] Users upload materials for video content, such as photos, audio, and music, from their devices to the server. The input is the digital media files (photos, audio, and music) uploaded by the user, and the output is a message confirming the completion of the upload. The device sends the provided materials to the server, which then stores them.
[2295] Step 6:
[2296] The server analyzes the uploaded material and proposes a style for the video content based on it. The input of this step is the uploaded material and emotional data, and the output is a recommended video style. The server analyzes the user's material and emotional state and recommends an appropriate style (e.g., an emotional anime style).
[2297] Step 7:
[2298] The user selects a suggested video style and then selects further options. The input is the style suggestion from the server and the user's selection, and the output is the selected style and options. The user makes their selection through the interface, and that information is sent to the server.
[2299] Step 8:
[2300] The server generates video content using a generative AI model based on the selected style and options. The input for this step is the user's selection information and emotional data, which are converted into specific prompt sentences. The output is the generated video content. The server sends the following prompt sentence to the generative AI model:
[2301] Create a video in anime style about: My graduation ceremony was one of the most memorable days of my life. I walked across the stage and received my diploma, surrounded by friends and family. There were moments of joy, laughter, and a bit of nostalgia as we looked back at our journey through university life.
[2302] Step 9:
[2303] The user can check the generated video on the preview page and request corrections if necessary. The input is the generated video content, and the output is the user's feedback or correction requests. The server displays the generated video to the user and receives any corrections that need to be made.
[2304] Step 10:
[2305] The server completes the revised video. The input of this step is the user's feedback, and the output is the completed video content. The server regenerates the video based on the user's feedback and provides the final version.
[2306] Step 11:
[2307] The user sets the visibility of the completed video content. The input is the visibility selection, and the output is a confirmation message that the settings have been completed. The server centrally makes the video available to other users based on the visibility (private, limited, public) entered.
[2308] Step 12:
[2309] The server monitors viewers' reactions to video content and collects emotional information. The input is viewer reaction data, and the output is analyzed emotional information. The server monitors viewers' reactions and comments and records them as reference information for the next video generation.
[2310] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2311] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2312] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2313] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2314] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2315] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2316] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2317] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2318] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2319] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinat...
Claims
1. A way for users to input information about events and turning points in their lives; A generating device generates video content based on input information; a display device for displaying the generated video content; A means for setting the scope of disclosure; A means for publishing the video content to other users according to the publishing scope; A system including:
2. a means for uploading multiple user-provided materials; A means of analyzing and storing the uploaded materials; The system of claim 1 further comprising:
3. A means for selecting different images when generating video content; means for automatically changing the style of the video content based on different selections; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A