System

The system addresses the limitations of conventional photo album apps by generating and displaying images based on user preferences, improving usability through voice recognition and continuous optimization.

JP2026015054APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116528
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional photo album apps only display photos that the user has actually taken, lacking the ability to automatically generate and add appropriate photos based on user preferences, and require complex operating procedures, making it difficult to create 'happy false memories' and posing usability challenges.

Method used

A system that inputs user taste and preference data to a server, generates images based on this data, transmits them to the user's terminal, displays the images, collects operation history, and updates the generation algorithm to improve usability by interpreting user requests through voice recognition and saving images in local storage.

Benefits of technology

Enables users to easily and intuitively generate and view high-quality images that match their preferences, creating 'happy false memories' and optimizing image generation based on user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015054000001_ABST
    Figure 2026015054000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting hobby and taste data of a user; means for transmitting data of the user to a server; means for generating an image based on the received data; means for transmitting the generated image to a terminal of the user; means for displaying the transmitted image on the terminal; means for collecting an operation history of the user; and means for updating a generation algorithm based on the operation history.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional photo album apps only display photos that the user has actually taken, providing a limited experience. Furthermore, they lack the ability to automatically generate and add appropriate photos based on the user's preferences, making it difficult to create "happy false memories" that the user has not experienced. Furthermore, they require complex operating procedures, posing usability challenges. The purpose of this invention is to solve these problems and provide users with a new photo experience while making user operations simple and intuitive. [Means for solving the problem]

[0005] The present invention includes a means for inputting a user's taste and preference data and transmitting it to a server. The server generates an image based on the received data and transmits the generated image to the user's terminal. The system also includes a means for displaying the image transmitted to the terminal and collecting the user's operation history. The system further includes a means for updating the generation algorithm based on the operation history. As a result, new photos that match the user's tastes are automatically generated and displayed. Usability is significantly improved by adding a means for interpreting a user's request using voice recognition and generating an image, and a means for saving the generated image in local storage and immediately displaying it to the user.

[0006] "User" refers to the individual who operates the system and generates or views images.

[0007] "Hobby and preference data" refers to information related to themes and genres that users prefer.

[0008] "Server" refers to a remote computing device that receives data sent by a user and performs processing and image generation.

[0009] "Terminal" means a device used by a user that has the function of displaying images sent from a server.

[0010] "Image generation" refers to the process of creating new images using AI technology.

[0011] "Transmission" refers to the act of sending data or generated images to another device (server, terminal, etc.).

[0012] "Operation history" refers to a record of a series of operations performed by a user on a system.

[0013] A "generative algorithm" refers to a computational method for creating new images based on input data.

[0014] "Speech recognition" refers to the technology that understands user requests input by voice and converts them into text.

[0015] "Local storage" refers to a storage device within a terminal, and is a location for saving received image data. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention relates to an album application for providing "fake happy memories" that are not taken by the user, and is implemented as follows.

[0038] User Data Collection

[0039] User: When launching the album app for the first time, the user enters their hobbies and interests on the initial setup screen, including specific information such as "I like traveling" and "I take lots of photos of my pets."

[0040] Terminal: Processes the data entered by the user and sends it to the server, where it is converted into an appropriate format and temporarily stored.

[0041] Server: Analyzes the received user data and stores it in a database, which provides the necessary information for later image generation.

[0042] Generate initial album

[0043] Server: Generates images using AI technology based on the user's hobbies and preferences. For example, if a user enters data such as "I like natural scenery," the AI ​​will generate an appropriate photo of a natural scenery.

[0044] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0045] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[0046] User interaction and feedback

[0047] User: Browse albums with touch or use voice input to request new images, for example, "Show me my new mountain climbing photos."

[0048] On your device: Detects touch input and displays details of selected photos, and sends voice input to a speech recognition engine for conversion to text.

[0049] Server: Receives and analyzes voice requests to extract new image generation requirements. For example, following the instruction "mountain climbing photos," it generates new mountain climbing photos.

[0050] Server: Sends the generated new image to the user device.

[0051] On the device: New images are received and saved to local storage, and are immediately displayed to the user.

[0052] Continuous learning and optimization

[0053] Server: Collects user activity history and records which photos are most frequently viewed and which requests are most frequently made.

[0054] Server: Using the collected operation history, the generation algorithm is updated and optimized, so that subsequent image generation will better match the user's preferences.

[0055] Specific examples

[0056] First-time setup

[0057] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take lots of photos of my pets."

[0058] Terminal: Sends input data to the server.

[0059] Server: Analyzes the received data and generates travel photos and pet photos.

[0060] Terminal: Displays the initially generated photo.

[0061] Daily use

[0062] User: While browsing an album, say "Show me new beach photos."

[0063] Device: Sends voice requests to the server.

[0064] Server: Generates beach photos and sends them to the user device.

[0065] Device: Generated beach photos are saved to local storage and displayed immediately.

[0066] In this way, users can enjoy the experience of gaining new, happy "false memories." By using this system, users can easily enjoy a variety of photos that suit their preferences.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] User: Launch the Album app. On the initial setup screen, enter your name and hobbies (e.g., "Travel," "Nature Scenery," "Pet Photos").

[0070] Step 2:

[0071] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[0072] Step 3:

[0073] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0074] Step 4:

[0075] Server: Based on the prepared parameters, images for the initial album are generated using a generative AI (e.g., GAN or VQ-VAE-2). The generated images are saved as temporary files.

[0076] Step 5:

[0077] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[0078] Step 6:

[0079] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[0080] Step 7:

[0081] User: Browse albums with touch input, select photos to view and navigate to details, and use voice input to request new photos (e.g., "Show me photos of the beach").

[0082] Step 8:

[0083] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, uses a speech recognition engine to convert speech into text and send it to the server.

[0084] Step 9:

[0085] Server: Analyzes the voice request and extracts parameters for generating a new image. Based on the extracted parameters, a new image is generated using generative AI.

[0086] Step 10:

[0087] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[0088] Step 11:

[0089] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0090] Step 12:

[0091] Server: Continuously collects user activity history, recording which photos are most frequently viewed and what requests are most frequently made.

[0092] Step 13:

[0093] Server: Updates and optimizes the generation algorithm based on the collected operation history, so that subsequent image generation can be more tailored to the user's preferences.

[0094] By following these steps, users can easily and intuitively operate the app to create new "happy false memories."

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] In modern digital album applications, many users spend time and effort efficiently collecting photos that match their hobbies and preferences. However, there are limited ways for users to easily obtain high-quality images that meet their individual needs without having to go through the trouble of manually collecting and selecting content. Furthermore, there are no systems that automatically optimize generated images using user operation history. Furthermore, the ability to interpret user requests using voice recognition and respond immediately is also insufficient. Therefore, the objective of this invention is to provide a system that provides "happy false memories" that match the user's preferences and are continuously optimized.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for inputting user preference information, means for transmitting the preference information to the server, means for generating images based on the received preference information, means for transmitting the generated images to the user's device, means for displaying the images transmitted to the device, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for interpreting the user's request using voice recognition and generating images, and means for saving the generated images in a local storage device and immediately displaying them to the user. This allows the user to easily and efficiently obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" optimized for their individual preferences.

[0100] "User" refers to an individual who uses the Album Application.

[0101] "Preference information" refers to data related to a user's hobbies and preferences that is entered on the initial setup screen, etc.

[0102] A "server" refers to a computing system that receives data from users and performs a series of processes such as analysis, image generation, and database management.

[0103] "Device" refers to the end-user devices used by a User, such as a smartphone, tablet, or PC.

[0104] "Means for generating images" refers to methods or functions for creating new images using a generative AI model based on received preference information.

[0105] "Voice recognition" refers to the technology that converts requests input by voice by the user into text data.

[0106] "Operation history" refers to data and logs generated by operations performed when a user uses the album application.

[0107] "Means for updating the generation algorithm" refers to a method for updating the training data of the AI ​​model based on collected operation history to improve the accuracy and efficiency of the image generation algorithm.

[0108] "Local storage device" refers to the data storage area that exists within the user's device, and specifically includes internal storage and external memory.

[0109] "Happy false memories" refer to fictional photos or image data that bring about a sense of happiness based on the user's preferences or requests, even though the user has not actually experienced them.

[0110] This invention relates to an album application that generates and provides "happy false memories" based on user preference information. This system involves a series of processes that collect user preference information, generate images using a generative AI model based on that information, and continuously optimize the images based on the user's operation history.

[0111] User Data Collection

[0112] When a user launches the album application for the first time, an initial setup screen appears. There, the user enters specific preferences, such as "I like traveling" or "I take lots of photos of my pet." The device converts this data into an appropriate format (e.g., JSON) and sends it to the server. The server analyzes the received data and stores it in a database.

[0113] Image generation

[0114] The server utilizes a generative AI model (e.g., DALL-E) to generate images based on the user's preferences. As a specific implementation example, a prompt such as "Based on the data that the user says 'I like traveling,' please generate photos of natural scenery at travel destinations" is input to the generative AI model. The server then assembles the generated images into a data packet and sends it to the user's device. The device then stores the received data in local storage and displays it in the album app.

[0115] User interaction and feedback

[0116] Users can browse the album using touch gestures and, if necessary, use voice input to request the generation of new images. For example, they could say, "Show me a new mountain climbing photo." The device receives the speech through its microphone and converts it into text using a speech recognition engine (e.g., Google Speech-to-Text). The text is sent to the server, which analyzes the request and generates a new prompt. For example, "Please generate a mountain climbing photo." The generated photo is sent back to the user's device, saved to local storage, and immediately displayed in the album app.

[0117] Continuous learning and optimization

[0118] The server collects user activity history, records which photos are most frequently viewed, and records which are most frequently requested, which is used to update the training dataset for the generative AI model and optimize the generation algorithm, so that subsequent image generation will be more in line with the user's preferences.

[0119] Specific examples

[0120] First-time setup

[0121] User: Launches the album app for the first time and enters preference information such as "I like traveling" and "I take lots of photos of my pets."

[0122] Terminal: Converts input data into JSON format and sends it to the server via an HTTP request.

[0123] Server: Analyzes the received data and stores it in a database. Prompts are input into the AI ​​model to generate images, such as travel photos and pet photos based on hobbies and preferences.

[0124] Device: The initially generated photos are saved to local storage and displayed in the album app.

[0125] Daily use

[0126] User: While browsing an album, make a voice request: "Show me new beach photos."

[0127] Device: Receives voice requests, converts them into text using a speech recognition engine, and sends the text to the server.

[0128] Server: Creates prompts to generate beach photos based on voice requests, inputs them into the generative AI model, and sends the generated beach photos to the user's device.

[0129] Device: Save the generated beach photos to local storage and display them instantly in the Album app.

[0130] The present invention allows users to easily obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" that are optimized to their individual preferences.

[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0132] Step 1: Collect user data

[0133] User: Launches the album app for the first time and enters hobby and preference data (e.g., "I like traveling" or "I take lots of photos of my pets") on the initial setup screen.

[0134] Terminal: Converts the data entered by the user into JSON format and temporarily stores it in local storage.

[0135] Device: Sends the saved data to the server as an HTTP request.

[0136] Server: Analyzes the received data and stores it in the user database. For analysis, the Python pandas library is used to format the data and extract the necessary fields.

[0137] Input: User-entered information about your interests and preferences.

[0138] Output: Cleaned and formatted database entries.

[0139] Step 2: Image generation

[0140] Server: Based on the user's preference information, the server sends a prompt to the generative AI model (e.g., DALL-E). It creates a specific prompt such as, "Based on the data that the user says 'I like natural scenery,' please generate a photo of a natural scenery."

[0141] Server: Receives the image data returned by the generative AI model and converts it into an appropriate format, such as JPEG.

[0142] Server: Builds the image data into a data packet and sends it to the user's device as an HTTP response.

[0143] Input: User preference information, prompts for the generative AI model.

[0144] Output: The generated image data.

[0145] Step 3: Album View

[0146] Terminal: Analyzes the received image data and saves it to a local storage device in a common image format such as JPEG.

[0147] On the device: Generate thumbnails of the images and display them in the album app. For example, use the Pillow library to generate thumbnails.

[0148] Input: Image data sent from the server.

[0149] Output: Image files saved to local storage, thumbnail display in app.

[0150] Step 4: User interaction and feedback

[0151] 4.1 Album browsing

[0152] User: Browse albums within the app using touch controls.

[0153] Device: Detects touch input and enlarges the image according to the touch position. The image enlargement process uses the built-in graphics library.

[0154] Input: Touch input data.

[0155] Output: The enlarged image.

[0156] 4.2 Voice Requests

[0157] User: Requests the generation of a new image via speech input (e.g., "Show me a new mountain climbing photo").

[0158] Device: Voice input is received via a microphone and converted into text using a speech recognition engine (e.g., Google Speech-to-Text).

[0159] Terminal: The converted text data is sent to the server as an HTTP request.

[0160] Input: The user's voice request.

[0161] Output: Data converted to text by the speech recognition engine and sent as an HTTP request.

[0162] 4.3 Creating and displaying new images

[0163] Server: Analyzes the voice request and inputs a new prompt to the generative AI model. For example, send a prompt such as "Generate a photo of mountain climbing."

[0164] Server: Sends the generated new image data to the user's terminal.

[0165] Device: New images received are saved to local storage and immediately displayed in the Album app.

[0166] Input: A prompt to the generative AI model via a voice request.

[0167] Output: New image data generated, new image displayed.

[0168] Step 5: Continue learning and optimization

[0169] Server: Collects user activity history, including browsing frequency, type of image requested, etc.

[0170] Server: Updates the training dataset for the generative AI model based on the collected operation history and optimizes the image generation algorithm. For updates, machine learning libraries such as Scikit-learn and TensorFlow are used.

[0171] Input: User operation history.

[0172] Output: An optimized generative AI model.

[0173] Through these steps, users can easily obtain and enjoy high-quality "happy false memories" based on their preferences.

[0174] (Application example 1)

[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] Conventional photo album applications create albums based on photos that users have actually taken, so unless users have experience taking photos, the albums are not complete. Also, it is difficult for users to easily collect photos that suit their preferences, which limits the user experience.

[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0178] In this invention, the server includes means for generating images using a generative AI model based on the user's hobby and preference data, means for sending the generated images to the user's device, and means for updating the generation algorithm using prompts based on the user's operation history. This allows the user to easily enjoy "happy false memories" that were not captured, and a variety of photos and videos that suit the user's preferences can be generated and provided in real time.

[0179] "User's hobby and preference data" is information entered by the user that represents the user's interests, concerns, favorite activities, and subjects.

[0180] "Server" means a computer system that receives, processes, stores, and generates data submitted by users.

[0181] A "generative AI model" is an algorithm that uses artificial intelligence to generate new images and videos based on user preferences and taste data.

[0182] The "means for generating images" refers to a mechanism that uses a generative AI model to create images based on the user's taste and preference data.

[0183] The "means for transmitting images to the user's terminal" is a communication protocol for transferring the generated image data to the device used by the user.

[0184] "Means for displaying images" refers to the functionality for visually displaying images generated on the user's device.

[0185] "User operation history" refers to records of touch operations, voice input, browsing history, etc., when a user uses an application.

[0186] A "prompt" is a textual instruction that instructs a generative AI model to generate a specific image or video.

[0187] "Speech recognition technology" is a technology that interprets and processes a user's verbal instructions by analyzing voice data and converting it into text data.

[0188] "Local storage" is a data storage area built into the user's device.

[0189] The present invention relates to a content distribution system for providing "happy false memories" that are not captured by the user, and is implemented as follows.

[0190] Device and Software Configuration

[0191] The hardware used is smartphones, smart glasses, and head-mounted displays (HMDs), such as Google Glass, North Focals, Oculus Rift, and HTC Vive. The software used is Google Cloud Speech-to-Text for voice recognition, generative AI models (e.g., OpenAI's DALL-E, Stable Diffusion) for image generation, and Firebase for the database.

[0192] User Data Collection

[0193] When users first launch the application, they enter their interests and preferences, which can include specific details like "I like to travel" or "I take lots of photos of my pets." This data is sent from their smartphone or device to the cloud and stored in a Firebase database.

[0194] Generate initial album

[0195] The server retrieves the user's interest data from the Firebase database and passes prompts to the generative AI model to generate an image. For example, if a user enters "I like natural scenery," the generative AI model receives the prompt "Generate a photo of a peaceful mountain hike with clear skies and lush greenery." The generated image is constructed as a data packet and sent to the user's device. The received data is stored in local storage on the device, and the image is displayed.

[0196] User Actions

[0197] Users can request the generation of new images using touch or voice input. For example, by saying, "Show me a new beach photo," the voice data is converted to text through Google Cloud Speech-to-Text and sent to the server. The server then provides the generative AI model with a new prompt, which generates a new photo. This process instantly sends the newly generated image to the user's device and displays it.

[0198] Continuous learning and optimization

[0199] The server collects user operation history, records which photos are viewed most frequently, and records which are frequently requested. Based on the collected operation history, the algorithm of the generative AI model is updated and optimized, so that future image generation will be more in line with the user's preferences.

[0200] Specific examples

[0201] For example, if a user makes a voice request such as "Show me new mountain climbing photos," the following prompt sentence is passed to the generative AI model:

[0202] "Generate a photo of a peaceful mountain hike with clear skies and lush greenery"

[0203] Through the above process, users can create and enjoy a variety of photos and videos that suit their tastes in real time.

[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0205] Step 1:

[0206] When a user launches the application for the first time, they enter their hobbies and preferences, such as "I like traveling" or "I take lots of photos of my pets." The entered data is sent from the smartphone or device to the server. The data is then temporarily processed on the user's device and converted into the appropriate format.

[0207] Step 2:

[0208] The server receives the interest and preference data sent by the user and stores it in the Firebase database. The input here is the interest and preference data, and the output is the stored data. The server analyzes the received data and prepares the information necessary for later image generation.

[0209] Step 3:

[0210] The server retrieves interest and preference data from the Firebase database and generates an image by passing a prompt to the generative AI model. For example, a specific prompt such as "Generate a photo of a peaceful mountain hike with clear skies and lush greenery" is provided to the generative AI model. The input is the interest and preference data and the prompt, and the output is the generated image.

[0211] Step 4:

[0212] The generated image is sent as a data packet from the server to the user's device. The input here is the generated image data, and the output is the image data received by the user's terminal. The server converts the image data into an appropriate format and transmits it via a communication protocol.

[0213] Step 5:

[0214] The device stores the image data received from the server in local storage and displays the image, where the input is the received image data and the output is the image visually displayed to the user. The device displays the stored image data in a user interface in an appropriate format.

[0215] Step 6:

[0216] The user requests the generation of a new image using touch or voice input. For example, they might say, "Show me a new beach photo." This voice data is picked up by the device's microphone and converted to text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data.

[0217] Step 7:

[0218] The device sends a voice request, converted into text, to the server, which analyzes the request and generates a new prompt. For example, a prompt such as "Generate a photo of a sunny beach with clear blue water" is passed to a generative AI model. The input is text data, and the output is the new prompt.

[0219] Step 8:

[0220] The server provides the new prompt to the generative AI model and generates a new image. The generated image is then sent back to the user's device. The input is the new prompt and the generative AI model, and the output is the newly generated image.

[0221] Step 9:

[0222] The device saves the newly received image to local storage and immediately displays it to the user. Here, the input is the newly received image data and the output is the image displayed to the user, allowing the user to immediately view the new content they requested.

[0223] Step 10:

[0224] The server continuously collects user operation history, recording which photos are viewed most frequently and which requests are most frequently made. This data is saved as operation history. The input is the user operation data, and the output is the saved operation history.

[0225] Step 11:

[0226] The server updates and optimizes the generative AI model algorithm based on the collected operation history. For example, if there are many requests for a particular situation, the server strengthens the generative algorithm to accommodate that situation. The input is operation history data, and the output is an optimized generative AI model. This process improves image generation from the next time onwards to better match the user's preferences.

[0227] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0228] The present invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that are not actually taken, and is implemented as follows.

[0229] User Data Collection

[0230] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0231] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[0232] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0233] Introducing the Emotion Engine

[0234] Device: The device acquires the user's facial expressions and tone of voice, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server along with the user data.

[0235] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[0236] Generate initial album

[0237] Server: Generates images using AI technology based on the user's hobby, preference, and emotional data. For example, if a user is recognized as "liking natural scenery" and "relaxed," a photo of a serene natural landscape will be generated.

[0238] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0239] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[0240] User interaction and feedback

[0241] Users can browse albums with touch or use voice input to request new images (e.g., "Show me pictures of the beach").

[0242] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[0243] Server: Analyzes the voice request and extracts parameters for generating new images. For example, based on the instruction "Photo of mountain climbing," it generates an appropriate photo and incorporates emotional data.

[0244] Server: Sends the generated new image to the user device.

[0245] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0246] Continuous learning and optimization

[0247] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[0248] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[0249] Specific examples

[0250] First-time setup

[0251] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take a lot of photos of my pets."

[0252] Terminal: Sends input data and the user's initial emotion data to the server.

[0253] Server: Analyzes hobby and preference data and emotional data to generate appropriate travel and pet photos.

[0254] Terminal: Displays the initially generated photo.

[0255] Daily use

[0256] User: While browsing an album, say "Show me new beach photos."

[0257] Device: Sends voice requests and emotion data to the server.

[0258] Server: Generates and sends beach photos based on the user's request and emotion data.

[0259] Device: Generated beach photos are saved to local storage and displayed immediately.

[0260] This system allows users to easily and intuitively experience "happy false memories" that match a specific emotional state.

[0261] The processing flow will be explained below.

[0262] Step 1:

[0263] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0264] Step 2:

[0265] Terminal: Temporarily stores the data entered by the user and sends it to the server, where it is converted into the appropriate format.

[0266] Step 3:

[0267] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0268] Step 4:

[0269] Device: Captures the user's facial expressions and voice tone, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server.

[0270] Step 5:

[0271] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[0272] Step 6:

[0273] Server: Generates images for the initial album using AI technology based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," it generates photos of tranquil natural scenery.

[0274] Step 7:

[0275] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[0276] Step 8:

[0277] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[0278] Step 9:

[0279] User: Browse the album with touch input or use voice input to request the generation of new images, for example, "Show me new beach photos."

[0280] Step 10:

[0281] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[0282] Step 11:

[0283] Server: Receives voice requests and emotion data, analyzes them, and extracts new image generation parameters. For example, in response to a request for a "beach photo," it generates an appropriate beach photo, taking into account the user's emotion data.

[0284] Step 12:

[0285] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[0286] Step 13:

[0287] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0288] Step 14:

[0289] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[0290] Step 15:

[0291] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[0292] By following the above steps, users can easily and intuitively experience a "happy false memory" that matches their specific emotional state.

[0293] Example 2

[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0295] With the recent advancement of digital technology, users are increasingly seeking tools that allow them to freely enjoy their imaginations and memories. However, existing album applications only collect and manage memories that users have clearly experienced, lacking the ability to generate new memories and experiences. Furthermore, they lack the ability to generate and display content that takes into account the user's emotional state, making it difficult to provide an experience that meets the user's individual needs and emotions. Furthermore, with few systems offering intuitive operation through voice input or real-time feedback, users are often forced to use complex operations.

[0296] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting user interest and preference data, means for transmitting user data to the server, means for acquiring user emotion data, means for generating an image based on the received user data and emotion data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting user operation history, means for updating the generation algorithm based on the operation history and emotion data, means for interpreting the user's request using voice recognition and generating a new image, and means for saving the generated image in local storage and immediately displaying it to the user. This makes it possible to easily generate new memories and experiences according to the user's interest and preference and emotional state, and intuitively provide content suitable for each individual user.

[0297] "User" means an individual or legal entity that uses the System.

[0298] "Hobby and preference data" refers to information about the genres and preferences that a user is interested in.

[0299] "Emotional data" refers to information about a user's emotional state obtained from facial expressions, tone of voice, etc.

[0300] "Server" refers to a computer system that processes, stores, and analyzes data submitted by users.

[0301] "Terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user.

[0302] "Received User Data" refers to information about a User that the Server receives from a Terminal.

[0303] "Image generation" refers to the process of creating new images using AI technology.

[0304] "Display means" refers to the method or technology that allows the generated image to be visually viewed on the user's device.

[0305] "Operation history" refers to the record of various operations performed by a user using the system.

[0306] "Generation algorithm" refers to the calculation procedure or model for generating an image based on received data.

[0307] A "voice recognition engine" refers to software that analyzes a user's voice and converts it into text.

[0308] "Local storage" refers to the data storage area within the user's device.

[0309] This invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that they have not taken. This system provides users with newly generated images using AI technology based on the user's taste data and emotion data.

[0310] Hardware and software used

[0311] Device: A device that is directly operated by the user, such as a smartphone, tablet, or personal computer, is used. The device is equipped with a camera and microphone, which are used to capture the user's facial expressions and voice.

[0312] Server: Use a high-performance cloud server or data center. The server performs central processing such as data storage, analysis, and image generation. Specific examples include servers provided by major cloud service providers (e.g., Amazon Web Services, Google Cloud Platform).

[0313] Software: The main software components of this system include an emotion engine, a speech recognition engine, and a generative AI model. We use OpenCV for the emotion engine and Google Cloud Speech-to-Text for the speech recognition engine. We use generative AI models (e.g., DALL·E, Stable Diffusion) for image generation.

[0314] Data flow and processing

[0315] 1. User data input

[0316] When a user launches the album app for the first time, they enter their name and hobbies (e.g., "travel," "nature scenery," "pet photos," etc.). The device temporarily stores this data, converts it into an appropriate format, and sends it to the server.

[0317] 2. Analysis and saving of initial setting data

[0318] The server analyzes the received user data and stores it in a database. During this analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[0319] 3. Acquiring and sending emotion data

[0320] The device's camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This emotion data is sent to a server, where it is analyzed and saved.

[0321] 4. Initial album generation

[0322] The server generates images using a generative AI model based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," a photo of a calm natural scenery will be generated. An example of a prompt sentence is "relaxing natural scenery."

[0323] 5. Sending and displaying images

[0324] The server then assembles the generated images into a data packet and sends it to the user's device, which receives it, stores it in local storage, and displays it in the album app.

[0325] User interaction and feedback

[0326] Users can browse the album with touch gestures and use voice input to request the generation of new images (e.g., "Show me new beach photos.") The device receives the voice request, converts it into text using a speech recognition engine, and sends it to the server.

[0327] The server analyzes the voice request and extracts parameters for generating a new image. For example, a prompt such as "beach photo" is input into the generative AI model, which then generates an image incorporating the user's emotional data. The image is then sent to the device, stored in local storage, and instantly displayed to the user.

[0328] Continuous learning and optimization

[0329] The server continuously collects the user's operation history and emotional data, and updates and optimizes the generation algorithm based on this data, so that the next image generation will be more in line with the user's tastes, preferences, and emotional state.

[0330] As described above, the present invention is a system that provides "happy false memories" that take into account the individual preferences and emotions of the user, and can provide the user with a new experience in an intuitive and simple manner.

[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0332] Step 1: User Data Input

[0333] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0334] Input: User name, hobbies and preferences

[0335] Output: Formatted user data (e.g., JSON format)

[0336] Specific behavior: The user enters their name and hobbies and preferences into the text boxes and check boxes, then presses the submit button.

[0337] Step 2: Send data

[0338] Terminal: Temporarily stores the input data, converts it into an appropriate format (e.g., JSON format), and sends it to the server.

[0339] Input: User data (name, hobbies, preferences)

[0340] Output: User data sent to the server

[0341] Specific behavior: After format conversion, the data is sent to the server using an HTTP request.

[0342] Step 3: Analyze and save the initial setup data

[0343] Server: Analyzes the received user data and stores it in a database. During the analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[0344] Input: Received user data

[0345] Output: Parsed user data stored in a database

[0346] Specific operation: Apply data analysis algorithms to generate parameters based on the user's preferences. Store the analysis results in a database.

[0347] Step 4: Obtaining and sending emotion data

[0348] Device: The camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This data is sent to the server.

[0349] Input: User's facial expression data, voice tone data

[0350] Output: Emotion data sent to the server

[0351] Specific operation: The camera captures faces in real time and records voice input. These data are analyzed by the emotion engine, and emotion data is extracted and sent to the server.

[0352] Step 5: Generate the initial album

[0353] Server: Generates images using a generative AI model based on the user's taste and emotion data.

[0354] Input: Interest data, emotion data, prompt sentence (e.g., "Relaxing nature scenery")

[0355] Output: Generated image data

[0356] Specific operation: A prompt sentence and emotion data are input to the generative AI model, and the AI ​​generates an image. The generated image is converted into a data packet.

[0357] Step 6: Send and view images

[0358] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0359] Device: The received image data is stored in local storage and displayed to the user in the album app.

[0360] Input: Generated image data

[0361] Output: Image displayed in the album app

[0362] Specific behavior: Receive image data from the server, save it to local storage, and display the image in the app's user interface.

[0363] Step 7: Browse albums and request new images

[0364] User: Browse albums with touch and use voice input to request new images (e.g., "Show me new beach photos").

[0365] Input: Touch, voice requests

[0366] Output: New image request data

[0367] Specific actions: Swipe to change images or use the microphone to input voice commands.

[0368] Step 8: Sending and analyzing audio data

[0369] Terminal: Receives voice requests, converts them into text using a speech recognition engine, and sends them to the server.

[0370] Input: Voice request

[0371] Output: The request text sent to the server

[0372] Specific operation: Converts speech to text and sends the text data to the server.

[0373] Step 9: Generate and send a new image

[0374] Server: Analyzes the voice request, extracts parameters for generating a new image, generates the image using a generative AI model, and sends it to the user's device.

[0375] Input: Voice request text, emotion data

[0376] Output: The new image data generated.

[0377] Specific operation: Analyzes the voice request and generates a prompt (e.g., "New beach photo"). Generates an image using a generative AI model, converts it into a data packet, and sends it to the device.

[0378] Step 10: Displaying the new image

[0379] On the device: The new image is generated and saved to local storage, and is immediately displayed to the user.

[0380] Input: Newly generated image data

[0381] Output: The new image displayed in the Album app

[0382] Specific behavior: Receives image data from the server, saves it to local storage, and displays the image in the album app's user interface.

[0383] Step 11: Continue learning and optimization

[0384] Server: Continuously collects user operation history and emotion data, and updates and optimizes the generation algorithm.

[0385] Input: Operation history, emotion data

[0386] Output: Updated and optimized generation algorithm

[0387] Specific operation: Obtain operation history and emotion data from the database. Based on this data, the generation algorithm is retrained and optimized.

[0388] (Application example 2)

[0389] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0390] While existing virtual stores recommend products based on users' preferences, they do not use real-time emotional data to recommend products that reflect the user's current emotional state. As a result, the accuracy of product recommendations to users is low, limiting the improvement of the user experience. In addition, systems linked to voice recognition have not been fully utilized when users make specific requests, making them difficult to use intuitively.

[0391] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0392] In this invention, the server includes means for inputting user interest and preference data, means for transmitting the user data to the server, means for generating an image based on the received data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for acquiring and analyzing the user's emotional data in real time, means for recommending products based on the acquired emotional data, and means for displaying the products on the terminal. This enables personalized product recommendations based on the user's emotional state in real time, improving the user experience and enabling intuitive operation.

[0393] "Hobbies and Preference Data" is information that indicates a user's interests, favorite activities, and hobbies.

[0394] "Emotional data" is data that represents the emotional state of a user, obtained from the user's facial expression, tone of voice, etc.

[0395] An "image generation means" is an algorithm or software that creates an image based on the data received.

[0396] The "product recommendation means" is a system that suggests products suitable for the user based on emotional data and hobby and preference data.

[0397] A "server" is a remote computer or system that has functions such as receiving, analyzing, and storing data, generating images, and recommending products.

[0398] "Terminal" means a device that can be directly operated by a user, including a smartphone, smart glasses, tablet, etc.

[0399] "Operation history" refers to historical data of operations and actions taken by a user when using the system.

[0400] "Speech recognition" is a technology that converts a user's voice into text data and understands the instructions based on that.

[0401] "Real-time" refers to data acquisition, analysis, and response occurring immediately and without delay.

[0402] This invention is a virtual store system that allows users to acquire their own emotional data in real time using their smart devices and recommends personalized products based on that data. To implement this invention, the following means are required.

[0403] First, the user puts on the smart glasses and starts the application. At startup, the user enters their name and hobby / preference data (e.g., "outdoor goods" or "casual fashion"). This data is temporarily stored in the smart glasses and sent to the server.

[0404] Next, the server prepares parameters for generating the initial album based on the received user's tastes and preferences. It also acquires the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine. The analyzed emotion data is sent to the server and stored in a database.

[0405] The server combines the user's taste and preference data with real-time emotional data to generate an image using a generative AI model. For example, if the user is in a "relaxed" state and prefers "outdoor goods," an image of outdoor goods with a relaxed atmosphere will be generated. The generated image and product information are organized into a data packet and sent to the user's smart glasses.

[0406] The smart glasses display the received image data and product information. The user can use voice requests to instruct the server to generate additional images. The voice requests are converted into text by a speech recognition engine, and the data is sent to the server. The server analyzes the voice requests and makes appropriate product recommendations.

[0407] The software components of this application include:

[0408] OpenCV: Used to analyze user facial expressions.

[0409] EmotionEngine: A custom AI model that identifies user emotions in real time.

[0410] ProductRecommender: An algorithm that provides personalized product recommendations based on user sentiment data.

[0411] ARDisplay: An augmented reality library that displays products virtually in the field of view of smart glasses.

[0412] As a concrete example, when a user enters a shopping mall and shows a relaxed expression, the emotion engine recognizes the state as "relaxed." Based on this, the server recommends "casual fashion items" and displays them on the smart glasses' display. When the user makes a voice request such as "Show me a new casual shirt," the voice is converted into text data and sent to the server. Based on this request, the server generates an image of an appropriate shirt and sends it back to the smart glasses.

[0413] Example prompt sentence:

[0414] Prompt: Recommend products suitable for the user when they are relaxed. The user is a man in his 30s who likes outdoor gear.

[0415] Example output: Recommendations include "Casual T-shirt", "Outdoor backpack", and "Sports sunglasses".

[0416] This enables personalized product suggestions based on user sentiment in real time, improving the user experience and providing an intuitive and convenient shopping experience.

[0417] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0418] Step 1:

[0419] Input of user's hobby and preference data

[0420] How it works: The user puts on the smart glasses and launches the application. On first launch, the user enters their name and preferences (e.g., "outdoor gear" or "casual fashion").

[0421] Input: User's basic information and hobbies and preferences

[0422] Data processing: The input data is temporarily stored in the smart glasses and converted into the required format.

[0423] Output: Formatted user data

[0424] Step 2:

[0425] Sending user data

[0426] How it works: The smart glasses send formatted user data to the server.

[0427] Input: Formatted user data

[0428] Data processing: None

[0429] Output: User data sent to the server

[0430] Step 3:

[0431] First, prepare to generate an album based on hobby and preference data.

[0432] Operation: The server analyzes the received user data and prepares parameters for generating the initial album based on the user's tastes and preferences.

[0433] Input: User data sent to the server

[0434] Data Calculation: Analyzing user data

[0435] Output: Parameters for initial album generation

[0436] Step 4:

[0437] Emotion data collection and analysis

[0438] How it works: The smart glasses capture the user's facial expressions and voice tone in real time, analyze these data with the emotion engine, and send the emotion data to the server.

[0439] Input: User facial expressions and voice tone

[0440] Data Computation: Emotion Analysis with Emotion Engine

[0441] Output: Emotion data

[0442] Step 5:

[0443] Emotion data storage and analysis

[0444] How it works: The server analyzes the received emotion data and stores it in a database.

[0445] Input: Emotion data

[0446] Data Computing: Sentiment Data Analysis and Storage

[0447] Output: Emotion data stored in a database

[0448] Step 6:

[0449] Image and product recommendation generation

[0450] How it works: By combining user preference data with real-time emotional data, a generative AI model is used to generate images and prepare product information.

[0451] Input: Parameters for initial album generation, emotion data

[0452] Data processing: Image generation using generative AI models, creation of product information

[0453] Output: Generated image data and product information

[0454] Step 7:

[0455] Submitting images and product information

[0456] Operation: The server assembles the generated image data and product information into a data packet and sends it to the user's smart glasses.

[0457] Input: Generated image data and product information

[0458] Data processing: Data packet construction

[0459] Output: Image data and product information sent to smart glasses

[0460] Step 8:

[0461] Display of images and product information

[0462] Operation: The smart glasses display the received image data and product information.

[0463] Input: Image data and product information sent to smart glasses

[0464] Data processing: None

[0465] Output: Image and product information displayed to the user

[0466] Step 9:

[0467] Processing user requests

[0468] How it works: A user uses a voice request to request additional images or product information. The smart glasses convert the voice request into text and send it to the server.

[0469] Input: User's voice request

[0470] Data calculation: Text conversion using a speech recognition engine

[0471] Output: Text request sent to the server

[0472] Step 10:

[0473] Generate and submit new images and products

[0474] How it works: The server analyzes the voice request, generates the appropriate images and products, and sends them to the smart glasses.

[0475] Input: The text request sent to the server

[0476] Data processing: Generative AI models generate new images and suggest products, and create data packets

[0477] Output: New image and product information sent to the smart glasses

[0478] Step 11:

[0479] New images and product displays

[0480] How it works: The smart glasses display the new images and product information they receive.

[0481] Input: New image and product information sent to the smart glasses

[0482] Data processing: None

[0483] Output: The new image and product information displayed to the user

[0484] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0485] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0486] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0487] [Second embodiment]

[0488] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0489] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0490] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0491] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0492] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0493] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0494] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0495] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0496] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0497] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0498] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0499] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0500] The present invention relates to an album application for providing "fake happy memories" that are not taken by the user, and is implemented as follows.

[0501] User Data Collection

[0502] User: When launching the album app for the first time, the user enters their hobbies and interests on the initial setup screen, including specific information such as "I like traveling" and "I take lots of photos of my pets."

[0503] Terminal: Processes the data entered by the user and sends it to the server, where it is converted into an appropriate format and temporarily stored.

[0504] Server: Analyzes the received user data and stores it in a database, which provides the necessary information for later image generation.

[0505] Generate initial album

[0506] Server: Generates images using AI technology based on the user's hobbies and preferences. For example, if a user enters data such as "I like natural scenery," the AI ​​will generate an appropriate photo of a natural scenery.

[0507] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0508] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[0509] User interaction and feedback

[0510] User: Browse albums with touch or use voice input to request new images, for example, "Show me my new mountain climbing photos."

[0511] On your device: Detects touch input and displays details of selected photos, and sends voice input to a speech recognition engine for conversion to text.

[0512] Server: Receives and analyzes voice requests to extract new image generation requirements. For example, following the instruction "mountain climbing photos," it generates new mountain climbing photos.

[0513] Server: Sends the generated new image to the user device.

[0514] On the device: New images are received and saved to local storage, and are immediately displayed to the user.

[0515] Continuous learning and optimization

[0516] Server: Collects user activity history and records which photos are most frequently viewed and which requests are most frequently made.

[0517] Server: Using the collected operation history, the generation algorithm is updated and optimized, so that subsequent image generation will better match the user's preferences.

[0518] Specific examples

[0519] First-time setup

[0520] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take lots of photos of my pets."

[0521] Terminal: Sends input data to the server.

[0522] Server: Analyzes the received data and generates travel photos and pet photos.

[0523] Terminal: Displays the initially generated photo.

[0524] Daily use

[0525] User: While browsing an album, say "Show me new beach photos."

[0526] Device: Sends voice requests to the server.

[0527] Server: Generates beach photos and sends them to the user device.

[0528] Device: Generated beach photos are saved to local storage and displayed immediately.

[0529] In this way, users can enjoy the experience of gaining new, happy "false memories." By using this system, users can easily enjoy a variety of photos that suit their preferences.

[0530] The processing flow will be explained below.

[0531] Step 1:

[0532] User: Launch the Album app. On the initial setup screen, enter your name and hobbies (e.g., "Travel," "Nature Scenery," "Pet Photos").

[0533] Step 2:

[0534] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[0535] Step 3:

[0536] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0537] Step 4:

[0538] Server: Based on the prepared parameters, images for the initial album are generated using a generative AI (e.g., GAN or VQ-VAE-2). The generated images are saved as temporary files.

[0539] Step 5:

[0540] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[0541] Step 6:

[0542] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[0543] Step 7:

[0544] User: Browse albums with touch input, select photos to view and navigate to details, and use voice input to request new photos (e.g., "Show me photos of the beach").

[0545] Step 8:

[0546] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, uses a speech recognition engine to convert speech into text and send it to the server.

[0547] Step 9:

[0548] Server: Analyzes the voice request and extracts parameters for generating a new image. Based on the extracted parameters, a new image is generated using generative AI.

[0549] Step 10:

[0550] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[0551] Step 11:

[0552] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0553] Step 12:

[0554] Server: Continuously collects user activity history, recording which photos are most frequently viewed and what requests are most frequently made.

[0555] Step 13:

[0556] Server: Updates and optimizes the generation algorithm based on the collected operation history, so that subsequent image generation can be more tailored to the user's preferences.

[0557] By following these steps, users can easily and intuitively operate the app to create new "happy false memories."

[0558] Example 1

[0559] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0560] In modern digital album applications, many users spend time and effort efficiently collecting photos that match their hobbies and preferences. However, there are limited ways for users to easily obtain high-quality images that meet their individual needs without having to go through the trouble of manually collecting and selecting content. Furthermore, there are no systems that automatically optimize generated images using user operation history. Furthermore, the ability to interpret user requests using voice recognition and respond immediately is also insufficient. Therefore, the objective of this invention is to provide a system that provides "happy false memories" that match the user's preferences and are continuously optimized.

[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0562] In this invention, the server includes means for inputting user preference information, means for transmitting the preference information to the server, means for generating images based on the received preference information, means for transmitting the generated images to the user's device, means for displaying the images transmitted to the device, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for interpreting the user's request using voice recognition and generating images, and means for saving the generated images in a local storage device and immediately displaying them to the user. This allows the user to easily and efficiently obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" optimized for their individual preferences.

[0563] "User" refers to an individual who uses the Album Application.

[0564] "Preference information" refers to data related to a user's hobbies and preferences that is entered on the initial setup screen, etc.

[0565] A "server" refers to a computing system that receives data from users and performs a series of processes such as analysis, image generation, and database management.

[0566] "Device" refers to the end-user devices used by a User, such as a smartphone, tablet, or PC.

[0567] "Means for generating images" refers to methods or functions for creating new images using a generative AI model based on received preference information.

[0568] "Voice recognition" refers to the technology that converts requests input by voice by the user into text data.

[0569] "Operation history" refers to data and logs generated by operations performed when a user uses the album application.

[0570] "Means for updating the generation algorithm" refers to a method for updating the training data of the AI ​​model based on collected operation history to improve the accuracy and efficiency of the image generation algorithm.

[0571] "Local storage device" refers to the data storage area that exists within the user's device, and specifically includes internal storage and external memory.

[0572] "Happy false memories" refer to fictional photos or image data that bring about a sense of happiness based on the user's preferences or requests, even though the user has not actually experienced them.

[0573] This invention relates to an album application that generates and provides "happy false memories" based on user preference information. This system involves a series of processes that collect user preference information, generate images using a generative AI model based on that information, and continuously optimize the images based on the user's operation history.

[0574] User Data Collection

[0575] When a user launches the album application for the first time, an initial setup screen appears. There, the user enters specific preferences, such as "I like traveling" or "I take lots of photos of my pet." The device converts this data into an appropriate format (e.g., JSON) and sends it to the server. The server analyzes the received data and stores it in a database.

[0576] Image generation

[0577] The server utilizes a generative AI model (e.g., DALL-E) to generate images based on the user's preferences. As a specific implementation example, a prompt such as "Based on the data that the user says 'I like traveling,' please generate photos of natural scenery at travel destinations" is input to the generative AI model. The server then assembles the generated images into a data packet and sends it to the user's device. The device then stores the received data in local storage and displays it in the album app.

[0578] User interaction and feedback

[0579] Users can browse the album using touch gestures and, if necessary, use voice input to request the generation of new images. For example, they could say, "Show me a new mountain climbing photo." The device receives the speech through its microphone and converts it into text using a speech recognition engine (e.g., Google Speech-to-Text). The text is sent to the server, which analyzes the request and generates a new prompt. For example, "Please generate a mountain climbing photo." The generated photo is sent back to the user's device, saved to local storage, and immediately displayed in the album app.

[0580] Continuous learning and optimization

[0581] The server collects user activity history, records which photos are most frequently viewed, and records which are most frequently requested, which is used to update the training dataset for the generative AI model and optimize the generation algorithm, so that subsequent image generation will be more in line with the user's preferences.

[0582] Specific examples

[0583] First-time setup

[0584] User: Launches the album app for the first time and enters preference information such as "I like traveling" and "I take lots of photos of my pets."

[0585] Terminal: Converts input data into JSON format and sends it to the server via an HTTP request.

[0586] Server: Analyzes the received data and stores it in a database. Prompts are input into the AI ​​model to generate images, such as travel photos and pet photos based on hobbies and preferences.

[0587] Device: The initially generated photos are saved to local storage and displayed in the album app.

[0588] Daily use

[0589] User: While browsing an album, make a voice request: "Show me new beach photos."

[0590] Device: Receives voice requests, converts them into text using a speech recognition engine, and sends the text to the server.

[0591] Server: Creates prompts to generate beach photos based on voice requests, inputs them into the generative AI model, and sends the generated beach photos to the user's device.

[0592] Device: Save the generated beach photos to local storage and display them instantly in the Album app.

[0593] The present invention allows users to easily obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" that are optimized to their individual preferences.

[0594] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0595] Step 1: Collect user data

[0596] User: Launches the album app for the first time and enters hobby and preference data (e.g., "I like traveling" or "I take lots of photos of my pets") on the initial setup screen.

[0597] Terminal: Converts the data entered by the user into JSON format and temporarily stores it in local storage.

[0598] Device: Sends the saved data to the server as an HTTP request.

[0599] Server: Analyzes the received data and stores it in the user database. For analysis, the Python pandas library is used to format the data and extract the necessary fields.

[0600] Input: User-entered information about your interests and preferences.

[0601] Output: Cleaned and formatted database entries.

[0602] Step 2: Image generation

[0603] Server: Based on the user's preference information, the server sends a prompt to the generative AI model (e.g., DALL-E). It creates a specific prompt such as, "Based on the data that the user says 'I like natural scenery,' please generate a photo of a natural scenery."

[0604] Server: Receives the image data returned by the generative AI model and converts it into an appropriate format, such as JPEG.

[0605] Server: Builds the image data into a data packet and sends it to the user's device as an HTTP response.

[0606] Input: User preference information, prompts for the generative AI model.

[0607] Output: The generated image data.

[0608] Step 3: Album View

[0609] Terminal: Analyzes the received image data and saves it to a local storage device in a common image format such as JPEG.

[0610] On the device: Generate thumbnails of the images and display them in the album app. For example, use the Pillow library to generate thumbnails.

[0611] Input: Image data sent from the server.

[0612] Output: Image files saved to local storage, thumbnail display in app.

[0613] Step 4: User interaction and feedback

[0614] 4.1 Album browsing

[0615] User: Browse albums within the app using touch controls.

[0616] Device: Detects touch input and enlarges the image according to the touch position. The image enlargement process uses the built-in graphics library.

[0617] Input: Touch input data.

[0618] Output: The enlarged image.

[0619] 4.2 Voice Requests

[0620] User: Requests the generation of a new image via speech input (e.g., "Show me a new mountain climbing photo").

[0621] Device: Voice input is received via a microphone and converted into text using a speech recognition engine (e.g., Google Speech-to-Text).

[0622] Terminal: The converted text data is sent to the server as an HTTP request.

[0623] Input: The user's voice request.

[0624] Output: Data converted to text by the speech recognition engine and sent as an HTTP request.

[0625] 4.3 Creating and displaying new images

[0626] Server: Analyzes the voice request and inputs a new prompt to the generative AI model. For example, send a prompt such as "Generate a photo of mountain climbing."

[0627] Server: Sends the generated new image data to the user's terminal.

[0628] Device: New images received are saved to local storage and immediately displayed in the Album app.

[0629] Input: A prompt to the generative AI model via a voice request.

[0630] Output: New image data generated, new image displayed.

[0631] Step 5: Continue learning and optimization

[0632] Server: Collects user activity history, including browsing frequency, type of image requested, etc.

[0633] Server: Updates the training dataset for the generative AI model based on the collected operation history and optimizes the image generation algorithm. For updates, machine learning libraries such as Scikit-learn and TensorFlow are used.

[0634] Input: User operation history.

[0635] Output: An optimized generative AI model.

[0636] Through these steps, users can easily obtain and enjoy high-quality "happy false memories" based on their preferences.

[0637] (Application example 1)

[0638] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0639] Conventional photo album applications create albums based on photos that users have actually taken, so unless users have experience taking photos, the albums are not complete. Also, it is difficult for users to easily collect photos that suit their preferences, which limits the user experience.

[0640] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0641] In this invention, the server includes means for generating images using a generative AI model based on the user's hobby and preference data, means for sending the generated images to the user's device, and means for updating the generation algorithm using prompts based on the user's operation history. This allows the user to easily enjoy "happy false memories" that were not captured, and a variety of photos and videos that suit the user's preferences can be generated and provided in real time.

[0642] "User's hobby and preference data" is information entered by the user that represents the user's interests, concerns, favorite activities, and subjects.

[0643] "Server" means a computer system that receives, processes, stores, and generates data submitted by users.

[0644] A "generative AI model" is an algorithm that uses artificial intelligence to generate new images and videos based on user preferences and taste data.

[0645] The "means for generating images" refers to a mechanism that uses a generative AI model to create images based on the user's taste and preference data.

[0646] The "means for transmitting images to the user's terminal" is a communication protocol for transferring the generated image data to the device used by the user.

[0647] "Means for displaying images" refers to the functionality for visually displaying images generated on the user's device.

[0648] "User operation history" refers to records of touch operations, voice input, browsing history, etc., when a user uses an application.

[0649] A "prompt" is a textual instruction that instructs a generative AI model to generate a specific image or video.

[0650] "Speech recognition technology" is a technology that interprets and processes a user's verbal instructions by analyzing voice data and converting it into text data.

[0651] "Local storage" is a data storage area built into the user's device.

[0652] The present invention relates to a content distribution system for providing "happy false memories" that are not captured by the user, and is implemented as follows.

[0653] Device and Software Configuration

[0654] The hardware used is smartphones, smart glasses, and head-mounted displays (HMDs), such as Google Glass, North Focals, Oculus Rift, and HTC Vive. The software used is Google Cloud Speech-to-Text for voice recognition, generative AI models (e.g., OpenAI's DALL-E, Stable Diffusion) for image generation, and Firebase for the database.

[0655] User Data Collection

[0656] When users first launch the application, they enter their interests and preferences, which can include specific details like "I like to travel" or "I take lots of photos of my pets." This data is sent from their smartphone or device to the cloud and stored in a Firebase database.

[0657] Generate initial album

[0658] The server retrieves the user's interest data from the Firebase database and passes prompts to the generative AI model to generate an image. For example, if a user enters "I like natural scenery," the generative AI model receives the prompt "Generate a photo of a peaceful mountain hike with clear skies and lush greenery." The generated image is constructed as a data packet and sent to the user's device. The received data is stored in local storage on the device, and the image is displayed.

[0659] User Actions

[0660] Users can request the generation of new images using touch or voice input. For example, by saying, "Show me a new beach photo," the voice data is converted to text through Google Cloud Speech-to-Text and sent to the server. The server then provides the generative AI model with a new prompt, which generates a new photo. This process instantly sends the newly generated image to the user's device and displays it.

[0661] Continuous learning and optimization

[0662] The server collects user operation history, records which photos are viewed most frequently, and records which are frequently requested. Based on the collected operation history, the algorithm of the generative AI model is updated and optimized, so that future image generation will be more in line with the user's preferences.

[0663] Specific examples

[0664] For example, if a user makes a voice request such as "Show me new mountain climbing photos," the following prompt sentence is passed to the generative AI model:

[0665] "Generate a photo of a peaceful mountain hike with clear skies and lush greenery"

[0666] Through the above process, users can create and enjoy a variety of photos and videos that suit their tastes in real time.

[0667] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0668] Step 1:

[0669] When a user launches the application for the first time, they enter their hobbies and preferences, such as "I like traveling" or "I take lots of photos of my pets." The entered data is sent from the smartphone or device to the server. The data is then temporarily processed on the user's device and converted into the appropriate format.

[0670] Step 2:

[0671] The server receives the interest and preference data sent by the user and stores it in the Firebase database. The input here is the interest and preference data, and the output is the stored data. The server analyzes the received data and prepares the information necessary for later image generation.

[0672] Step 3:

[0673] The server retrieves interest and preference data from the Firebase database and generates an image by passing a prompt to the generative AI model. For example, a specific prompt such as "Generate a photo of a peaceful mountain hike with clear skies and lush greenery" is provided to the generative AI model. The input is the interest and preference data and the prompt, and the output is the generated image.

[0674] Step 4:

[0675] The generated image is sent as a data packet from the server to the user's device. The input here is the generated image data, and the output is the image data received by the user's terminal. The server converts the image data into an appropriate format and transmits it via a communication protocol.

[0676] Step 5:

[0677] The device stores the image data received from the server in local storage and displays the image, where the input is the received image data and the output is the image visually displayed to the user. The device displays the stored image data in a user interface in an appropriate format.

[0678] Step 6:

[0679] The user requests the generation of a new image using touch or voice input. For example, they might say, "Show me a new beach photo." This voice data is picked up by the device's microphone and converted to text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data.

[0680] Step 7:

[0681] The device sends a voice request, converted into text, to the server, which analyzes the request and generates a new prompt. For example, a prompt such as "Generate a photo of a sunny beach with clear blue water" is passed to a generative AI model. The input is text data, and the output is the new prompt.

[0682] Step 8:

[0683] The server provides the new prompt to the generative AI model and generates a new image. The generated image is then sent back to the user's device. The input is the new prompt and the generative AI model, and the output is the newly generated image.

[0684] Step 9:

[0685] The device saves the newly received image to local storage and immediately displays it to the user. Here, the input is the newly received image data and the output is the image displayed to the user, allowing the user to immediately view the new content they requested.

[0686] Step 10:

[0687] The server continuously collects user operation history, recording which photos are viewed most frequently and which requests are most frequently made. This data is saved as operation history. The input is the user operation data, and the output is the saved operation history.

[0688] Step 11:

[0689] The server updates and optimizes the generative AI model algorithm based on the collected operation history. For example, if there are many requests for a particular situation, the server strengthens the generative algorithm to accommodate that situation. The input is operation history data, and the output is an optimized generative AI model. This process improves image generation from the next time onwards to better match the user's preferences.

[0690] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0691] The present invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that are not actually taken, and is implemented as follows.

[0692] User Data Collection

[0693] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0694] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[0695] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0696] Introducing the Emotion Engine

[0697] Device: The device acquires the user's facial expressions and tone of voice, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server along with the user data.

[0698] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[0699] Generate initial album

[0700] Server: Generates images using AI technology based on the user's hobby, preference, and emotional data. For example, if a user is recognized as "liking natural scenery" and "relaxed," a photo of a serene natural landscape will be generated.

[0701] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0702] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[0703] User interaction and feedback

[0704] Users can browse albums with touch or use voice input to request new images (e.g., "Show me pictures of the beach").

[0705] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[0706] Server: Analyzes the voice request and extracts parameters for generating new images. For example, based on the instruction "Photo of mountain climbing," it generates an appropriate photo and incorporates emotional data.

[0707] Server: Sends the generated new image to the user device.

[0708] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0709] Continuous learning and optimization

[0710] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[0711] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[0712] Specific examples

[0713] First-time setup

[0714] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take a lot of photos of my pets."

[0715] Terminal: Sends input data and the user's initial emotion data to the server.

[0716] Server: Analyzes hobby and preference data and emotional data to generate appropriate travel and pet photos.

[0717] Terminal: Displays the initially generated photo.

[0718] Daily use

[0719] User: While browsing an album, say "Show me new beach photos."

[0720] Device: Sends voice requests and emotion data to the server.

[0721] Server: Generates and sends beach photos based on the user's request and emotion data.

[0722] Device: Generated beach photos are saved to local storage and displayed immediately.

[0723] This system allows users to easily and intuitively experience "happy false memories" that match a specific emotional state.

[0724] The processing flow will be explained below.

[0725] Step 1:

[0726] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0727] Step 2:

[0728] Terminal: Temporarily stores the data entered by the user and sends it to the server, where it is converted into the appropriate format.

[0729] Step 3:

[0730] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[0731] Step 4:

[0732] Device: Captures the user's facial expressions and voice tone, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server.

[0733] Step 5:

[0734] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[0735] Step 6:

[0736] Server: Generates images for the initial album using AI technology based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," it generates photos of tranquil natural scenery.

[0737] Step 7:

[0738] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[0739] Step 8:

[0740] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[0741] Step 9:

[0742] User: Browse the album with touch input or use voice input to request the generation of new images, for example, "Show me new beach photos."

[0743] Step 10:

[0744] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[0745] Step 11:

[0746] Server: Receives voice requests and emotion data, analyzes them, and extracts new image generation parameters. For example, in response to a request for a "beach photo," it generates an appropriate beach photo, taking into account the user's emotion data.

[0747] Step 12:

[0748] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[0749] Step 13:

[0750] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[0751] Step 14:

[0752] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[0753] Step 15:

[0754] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[0755] By following the above steps, users can easily and intuitively experience a "happy false memory" that matches their specific emotional state.

[0756] Example 2

[0757] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0758] With the recent advancement of digital technology, users are increasingly seeking tools that allow them to freely enjoy their imaginations and memories. However, existing album applications only collect and manage memories that users have clearly experienced, lacking the ability to generate new memories and experiences. Furthermore, they lack the ability to generate and display content that takes into account the user's emotional state, making it difficult to provide an experience that meets the user's individual needs and emotions. Furthermore, with few systems offering intuitive operation through voice input or real-time feedback, users are often forced to use complex operations.

[0759] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting user interest and preference data, means for transmitting user data to the server, means for acquiring user emotion data, means for generating an image based on the received user data and emotion data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting user operation history, means for updating the generation algorithm based on the operation history and emotion data, means for interpreting the user's request using voice recognition and generating a new image, and means for saving the generated image in local storage and immediately displaying it to the user. This makes it possible to easily generate new memories and experiences according to the user's interest and preference and emotional state, and intuitively provide content suitable for each individual user.

[0760] "User" means an individual or legal entity that uses the System.

[0761] "Hobby and preference data" refers to information about the genres and preferences that a user is interested in.

[0762] "Emotional data" refers to information about a user's emotional state obtained from facial expressions, tone of voice, etc.

[0763] "Server" refers to a computer system that processes, stores, and analyzes data submitted by users.

[0764] "Terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user.

[0765] "Received User Data" refers to information about a User that the Server receives from a Terminal.

[0766] "Image generation" refers to the process of creating new images using AI technology.

[0767] "Display means" refers to the method or technology that allows the generated image to be visually viewed on the user's device.

[0768] "Operation history" refers to the record of various operations performed by a user using the system.

[0769] "Generation algorithm" refers to the calculation procedure or model for generating an image based on received data.

[0770] A "voice recognition engine" refers to software that analyzes a user's voice and converts it into text.

[0771] "Local storage" refers to the data storage area within the user's device.

[0772] This invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that they have not taken. This system provides users with newly generated images using AI technology based on the user's taste data and emotion data.

[0773] Hardware and software used

[0774] Device: A device that is directly operated by the user, such as a smartphone, tablet, or personal computer, is used. The device is equipped with a camera and microphone, which are used to capture the user's facial expressions and voice.

[0775] Server: Use a high-performance cloud server or data center. The server performs central processing such as data storage, analysis, and image generation. Specific examples include servers provided by major cloud service providers (e.g., Amazon Web Services, Google Cloud Platform).

[0776] Software: The main software components of this system include an emotion engine, a speech recognition engine, and a generative AI model. We use OpenCV for the emotion engine and Google Cloud Speech-to-Text for the speech recognition engine. We use generative AI models (e.g., DALL·E, Stable Diffusion) for image generation.

[0777] Data flow and processing

[0778] 1. User data input

[0779] When a user launches the album app for the first time, they enter their name and hobbies (e.g., "travel," "nature scenery," "pet photos," etc.). The device temporarily stores this data, converts it into an appropriate format, and sends it to the server.

[0780] 2. Analysis and saving of initial setting data

[0781] The server analyzes the received user data and stores it in a database. During this analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[0782] 3. Acquiring and sending emotion data

[0783] The device's camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This emotion data is sent to a server, where it is analyzed and saved.

[0784] 4. Initial album generation

[0785] The server generates images using a generative AI model based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," a photo of a calm natural scenery will be generated. An example of a prompt sentence is "relaxing natural scenery."

[0786] 5. Sending and displaying images

[0787] The server then assembles the generated images into a data packet and sends it to the user's device, which receives it, stores it in local storage, and displays it in the album app.

[0788] User interaction and feedback

[0789] Users can browse the album with touch gestures and use voice input to request the generation of new images (e.g., "Show me new beach photos.") The device receives the voice request, converts it into text using a speech recognition engine, and sends it to the server.

[0790] The server analyzes the voice request and extracts parameters for generating a new image. For example, a prompt such as "beach photo" is input into the generative AI model, which then generates an image incorporating the user's emotional data. The image is then sent to the device, stored in local storage, and instantly displayed to the user.

[0791] Continuous learning and optimization

[0792] The server continuously collects the user's operation history and emotional data, and updates and optimizes the generation algorithm based on this data, so that the next image generation will be more in line with the user's tastes, preferences, and emotional state.

[0793] As described above, the present invention is a system that provides "happy false memories" that take into account the individual preferences and emotions of the user, and can provide the user with a new experience in an intuitive and simple manner.

[0794] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0795] Step 1: User Data Input

[0796] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[0797] Input: User name, hobbies and preferences

[0798] Output: Formatted user data (e.g., JSON format)

[0799] Specific behavior: The user enters their name and hobbies and preferences into the text boxes and check boxes, then presses the submit button.

[0800] Step 2: Send data

[0801] Terminal: Temporarily stores the input data, converts it into an appropriate format (e.g., JSON format), and sends it to the server.

[0802] Input: User data (name, hobbies, preferences)

[0803] Output: User data sent to the server

[0804] Specific behavior: After format conversion, the data is sent to the server using an HTTP request.

[0805] Step 3: Analyze and save the initial setup data

[0806] Server: Analyzes the received user data and stores it in a database. During the analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[0807] Input: Received user data

[0808] Output: Parsed user data stored in a database

[0809] Specific operation: Apply data analysis algorithms to generate parameters based on the user's preferences. Store the analysis results in a database.

[0810] Step 4: Obtaining and sending emotion data

[0811] Device: The camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This data is sent to the server.

[0812] Input: User's facial expression data, voice tone data

[0813] Output: Emotion data sent to the server

[0814] Specific operation: The camera captures faces in real time and records voice input. These data are analyzed by the emotion engine, and emotion data is extracted and sent to the server.

[0815] Step 5: Generate the initial album

[0816] Server: Generates images using a generative AI model based on the user's taste and emotion data.

[0817] Input: Interest data, emotion data, prompt sentence (e.g., "Relaxing nature scenery")

[0818] Output: Generated image data

[0819] Specific operation: A prompt sentence and emotion data are input to the generative AI model, and the AI ​​generates an image. The generated image is converted into a data packet.

[0820] Step 6: Send and view images

[0821] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0822] Device: The received image data is stored in local storage and displayed to the user in the album app.

[0823] Input: Generated image data

[0824] Output: Image displayed in the album app

[0825] Specific behavior: Receive image data from the server, save it to local storage, and display the image in the app's user interface.

[0826] Step 7: Browse albums and request new images

[0827] User: Browse albums with touch and use voice input to request new images (e.g., "Show me new beach photos").

[0828] Input: Touch, voice requests

[0829] Output: New image request data

[0830] Specific actions: Swipe to change images or use the microphone to input voice commands.

[0831] Step 8: Sending and analyzing audio data

[0832] Terminal: Receives voice requests, converts them into text using a speech recognition engine, and sends them to the server.

[0833] Input: Voice request

[0834] Output: The request text sent to the server

[0835] Specific operation: Converts speech to text and sends the text data to the server.

[0836] Step 9: Generate and send a new image

[0837] Server: Analyzes the voice request, extracts parameters for generating a new image, generates the image using a generative AI model, and sends it to the user's device.

[0838] Input: Voice request text, emotion data

[0839] Output: The new image data generated.

[0840] Specific operation: Analyzes the voice request and generates a prompt (e.g., "New beach photo"). Generates an image using a generative AI model, converts it into a data packet, and sends it to the device.

[0841] Step 10: Displaying the new image

[0842] On the device: The new image is generated and saved to local storage, and is immediately displayed to the user.

[0843] Input: Newly generated image data

[0844] Output: The new image displayed in the Album app

[0845] Specific behavior: Receives image data from the server, saves it to local storage, and displays the image in the album app's user interface.

[0846] Step 11: Continue learning and optimization

[0847] Server: Continuously collects user operation history and emotion data, and updates and optimizes the generation algorithm.

[0848] Input: Operation history, emotion data

[0849] Output: Updated and optimized generation algorithm

[0850] Specific operation: Obtain operation history and emotion data from the database. Based on this data, the generation algorithm is retrained and optimized.

[0851] (Application example 2)

[0852] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0853] While existing virtual stores recommend products based on users' preferences, they do not use real-time emotional data to recommend products that reflect the user's current emotional state. As a result, the accuracy of product recommendations to users is low, limiting the improvement of the user experience. In addition, systems linked to voice recognition have not been fully utilized when users make specific requests, making them difficult to use intuitively.

[0854] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0855] In this invention, the server includes means for inputting user interest and preference data, means for transmitting the user data to the server, means for generating an image based on the received data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for acquiring and analyzing the user's emotional data in real time, means for recommending products based on the acquired emotional data, and means for displaying the products on the terminal. This enables personalized product recommendations based on the user's emotional state in real time, improving the user experience and enabling intuitive operation.

[0856] "Hobbies and Preference Data" is information that indicates a user's interests, favorite activities, and hobbies.

[0857] "Emotional data" is data that represents the emotional state of a user, obtained from the user's facial expression, tone of voice, etc.

[0858] An "image generation means" is an algorithm or software that creates an image based on the data received.

[0859] The "product recommendation means" is a system that suggests products suitable for the user based on emotional data and hobby and preference data.

[0860] A "server" is a remote computer or system that has functions such as receiving, analyzing, and storing data, generating images, and recommending products.

[0861] "Terminal" means a device that can be directly operated by a user, including a smartphone, smart glasses, tablet, etc.

[0862] "Operation history" refers to historical data of operations and actions taken by a user when using the system.

[0863] "Speech recognition" is a technology that converts a user's voice into text data and understands the instructions based on that.

[0864] "Real-time" refers to data acquisition, analysis, and response occurring immediately and without delay.

[0865] This invention is a virtual store system that allows users to acquire their own emotional data in real time using their smart devices and recommends personalized products based on that data. To implement this invention, the following means are required.

[0866] First, the user puts on the smart glasses and starts the application. At startup, the user enters their name and hobby / preference data (e.g., "outdoor goods" or "casual fashion"). This data is temporarily stored in the smart glasses and sent to the server.

[0867] Next, the server prepares parameters for generating the initial album based on the received user's tastes and preferences. It also acquires the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine. The analyzed emotion data is sent to the server and stored in a database.

[0868] The server combines the user's taste and preference data with real-time emotional data to generate an image using a generative AI model. For example, if the user is in a "relaxed" state and prefers "outdoor goods," an image of outdoor goods with a relaxed atmosphere will be generated. The generated image and product information are organized into a data packet and sent to the user's smart glasses.

[0869] The smart glasses display the received image data and product information. The user can use voice requests to instruct the server to generate additional images. The voice requests are converted into text by a speech recognition engine, and the data is sent to the server. The server analyzes the voice requests and makes appropriate product recommendations.

[0870] The software components of this application include:

[0871] OpenCV: Used to analyze user facial expressions.

[0872] EmotionEngine: A custom AI model that identifies user emotions in real time.

[0873] ProductRecommender: An algorithm that provides personalized product recommendations based on user sentiment data.

[0874] ARDisplay: An augmented reality library that displays products virtually in the field of view of smart glasses.

[0875] As a concrete example, when a user enters a shopping mall and shows a relaxed expression, the emotion engine recognizes the state as "relaxed." Based on this, the server recommends "casual fashion items" and displays them on the smart glasses' display. When the user makes a voice request such as "Show me a new casual shirt," the voice is converted into text data and sent to the server. Based on this request, the server generates an image of an appropriate shirt and sends it back to the smart glasses.

[0876] Example prompt sentence:

[0877] Prompt: Recommend products suitable for the user when they are relaxed. The user is a man in his 30s who likes outdoor gear.

[0878] Example output: Recommendations include "Casual T-shirt", "Outdoor backpack", and "Sports sunglasses".

[0879] This enables personalized product suggestions based on user sentiment in real time, improving the user experience and providing an intuitive and convenient shopping experience.

[0880] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0881] Step 1:

[0882] Input of user's hobby and preference data

[0883] How it works: The user puts on the smart glasses and launches the application. On first launch, the user enters their name and preferences (e.g., "outdoor gear" or "casual fashion").

[0884] Input: User's basic information and hobbies and preferences

[0885] Data processing: The input data is temporarily stored in the smart glasses and converted into the required format.

[0886] Output: Formatted user data

[0887] Step 2:

[0888] Sending user data

[0889] How it works: The smart glasses send formatted user data to the server.

[0890] Input: Formatted user data

[0891] Data processing: None

[0892] Output: User data sent to the server

[0893] Step 3:

[0894] First, prepare to generate an album based on hobby and preference data.

[0895] Operation: The server analyzes the received user data and prepares parameters for generating the initial album based on the user's tastes and preferences.

[0896] Input: User data sent to the server

[0897] Data Calculation: Analyzing user data

[0898] Output: Parameters for initial album generation

[0899] Step 4:

[0900] Emotion data collection and analysis

[0901] How it works: The smart glasses capture the user's facial expressions and voice tone in real time, analyze these data with the emotion engine, and send the emotion data to the server.

[0902] Input: User facial expressions and voice tone

[0903] Data Computation: Emotion Analysis with Emotion Engine

[0904] Output: Emotion data

[0905] Step 5:

[0906] Emotion data storage and analysis

[0907] How it works: The server analyzes the received emotion data and stores it in a database.

[0908] Input: Emotion data

[0909] Data Computing: Sentiment Data Analysis and Storage

[0910] Output: Emotion data stored in a database

[0911] Step 6:

[0912] Image and product recommendation generation

[0913] How it works: By combining user preference data with real-time emotional data, a generative AI model is used to generate images and prepare product information.

[0914] Input: Parameters for initial album generation, emotion data

[0915] Data processing: Image generation using generative AI models, creation of product information

[0916] Output: Generated image data and product information

[0917] Step 7:

[0918] Submitting images and product information

[0919] Operation: The server assembles the generated image data and product information into a data packet and sends it to the user's smart glasses.

[0920] Input: Generated image data and product information

[0921] Data processing: Data packet construction

[0922] Output: Image data and product information sent to smart glasses

[0923] Step 8:

[0924] Display of images and product information

[0925] Operation: The smart glasses display the received image data and product information.

[0926] Input: Image data and product information sent to smart glasses

[0927] Data processing: None

[0928] Output: Image and product information displayed to the user

[0929] Step 9:

[0930] Processing user requests

[0931] How it works: A user uses a voice request to request additional images or product information. The smart glasses convert the voice request into text and send it to the server.

[0932] Input: User's voice request

[0933] Data calculation: Text conversion using a speech recognition engine

[0934] Output: Text request sent to the server

[0935] Step 10:

[0936] Generate and submit new images and products

[0937] How it works: The server analyzes the voice request, generates the appropriate images and products, and sends them to the smart glasses.

[0938] Input: The text request sent to the server

[0939] Data processing: Generative AI models generate new images and suggest products, and create data packets

[0940] Output: New image and product information sent to the smart glasses

[0941] Step 11:

[0942] New images and product displays

[0943] How it works: The smart glasses display the new images and product information they receive.

[0944] Input: New image and product information sent to the smart glasses

[0945] Data processing: None

[0946] Output: The new image and product information displayed to the user

[0947] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0948] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0949] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0950] [Third embodiment]

[0951] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0952] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0953] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0954] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0955] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0956] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0957] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0958] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0959] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0960] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0961] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0962] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0963] The present invention relates to an album application for providing "fake happy memories" that are not taken by the user, and is implemented as follows.

[0964] User Data Collection

[0965] User: When launching the album app for the first time, the user enters their hobbies and interests on the initial setup screen, including specific information such as "I like traveling" and "I take lots of photos of my pets."

[0966] Terminal: Processes the data entered by the user and sends it to the server, where it is converted into an appropriate format and temporarily stored.

[0967] Server: Analyzes the received user data and stores it in a database, which provides the necessary information for later image generation.

[0968] Generate initial album

[0969] Server: Generates images using AI technology based on the user's hobbies and preferences. For example, if a user enters data such as "I like natural scenery," the AI ​​will generate an appropriate photo of a natural scenery.

[0970] Server: Assembles the generated image as a data packet and sends it to the user's device.

[0971] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[0972] User interaction and feedback

[0973] User: Browse albums with touch or use voice input to request new images, for example, "Show me my new mountain climbing photos."

[0974] On your device: Detects touch input and displays details of selected photos, and sends voice input to a speech recognition engine for conversion to text.

[0975] Server: Receives and analyzes voice requests to extract new image generation requirements. For example, following the instruction "mountain climbing photos," it generates new mountain climbing photos.

[0976] Server: Sends the generated new image to the user device.

[0977] On the device: New images are received and saved to local storage, and are immediately displayed to the user.

[0978] Continuous learning and optimization

[0979] Server: Collects user activity history and records which photos are most frequently viewed and which requests are most frequently made.

[0980] Server: Using the collected operation history, the generation algorithm is updated and optimized, so that subsequent image generation will better match the user's preferences.

[0981] Specific examples

[0982] First-time setup

[0983] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take lots of photos of my pets."

[0984] Terminal: Sends input data to the server.

[0985] Server: Analyzes the received data and generates travel photos and pet photos.

[0986] Terminal: Displays the initially generated photo.

[0987] Daily use

[0988] User: While browsing an album, say "Show me new beach photos."

[0989] Device: Sends voice requests to the server.

[0990] Server: Generates beach photos and sends them to the user device.

[0991] Device: Generated beach photos are saved to local storage and displayed immediately.

[0992] In this way, users can enjoy the experience of gaining new, happy "false memories." By using this system, users can easily enjoy a variety of photos that suit their preferences.

[0993] The processing flow will be explained below.

[0994] Step 1:

[0995] User: Launch the Album app. On the initial setup screen, enter your name and hobbies (e.g., "Travel," "Nature Scenery," "Pet Photos").

[0996] Step 2:

[0997] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[0998] Step 3:

[0999] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1000] Step 4:

[1001] Server: Based on the prepared parameters, images for the initial album are generated using a generative AI (e.g., GAN or VQ-VAE-2). The generated images are saved as temporary files.

[1002] Step 5:

[1003] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[1004] Step 6:

[1005] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[1006] Step 7:

[1007] User: Browse albums with touch input, select photos to view and navigate to details, and use voice input to request new photos (e.g., "Show me photos of the beach").

[1008] Step 8:

[1009] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, uses a speech recognition engine to convert speech into text and send it to the server.

[1010] Step 9:

[1011] Server: Analyzes the voice request and extracts parameters for generating a new image. Based on the extracted parameters, a new image is generated using generative AI.

[1012] Step 10:

[1013] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[1014] Step 11:

[1015] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1016] Step 12:

[1017] Server: Continuously collects user activity history, recording which photos are most frequently viewed and what requests are most frequently made.

[1018] Step 13:

[1019] Server: Updates and optimizes the generation algorithm based on the collected operation history, so that subsequent image generation can be more tailored to the user's preferences.

[1020] By following these steps, users can easily and intuitively operate the app to create new "happy false memories."

[1021] Example 1

[1022] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1023] In modern digital album applications, many users spend time and effort efficiently collecting photos that match their hobbies and preferences. However, there are limited ways for users to easily obtain high-quality images that meet their individual needs without having to go through the trouble of manually collecting and selecting content. Furthermore, there are no systems that automatically optimize generated images using user operation history. Furthermore, the ability to interpret user requests using voice recognition and respond immediately is also insufficient. Therefore, the objective of this invention is to provide a system that provides "happy false memories" that match the user's preferences and are continuously optimized.

[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1025] In this invention, the server includes means for inputting user preference information, means for transmitting the preference information to the server, means for generating images based on the received preference information, means for transmitting the generated images to the user's device, means for displaying the images transmitted to the device, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for interpreting the user's request using voice recognition and generating images, and means for saving the generated images in a local storage device and immediately displaying them to the user. This allows the user to easily and efficiently obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" optimized for their individual preferences.

[1026] "User" refers to an individual who uses the Album Application.

[1027] "Preference information" refers to data related to a user's hobbies and preferences that is entered on the initial setup screen, etc.

[1028] A "server" refers to a computing system that receives data from users and performs a series of processes such as analysis, image generation, and database management.

[1029] "Device" refers to the end-user devices used by a User, such as a smartphone, tablet, or PC.

[1030] "Means for generating images" refers to methods or functions for creating new images using a generative AI model based on received preference information.

[1031] "Voice recognition" refers to the technology that converts requests input by voice by the user into text data.

[1032] "Operation history" refers to data and logs generated by operations performed when a user uses the album application.

[1033] "Means for updating the generation algorithm" refers to a method for updating the training data of the AI ​​model based on collected operation history to improve the accuracy and efficiency of the image generation algorithm.

[1034] "Local storage device" refers to the data storage area that exists within the user's device, and specifically includes internal storage and external memory.

[1035] "Happy false memories" refer to fictional photos or image data that bring about a sense of happiness based on the user's preferences or requests, even though the user has not actually experienced them.

[1036] This invention relates to an album application that generates and provides "happy false memories" based on user preference information. This system involves a series of processes that collect user preference information, generate images using a generative AI model based on that information, and continuously optimize the images based on the user's operation history.

[1037] User Data Collection

[1038] When a user launches the album application for the first time, an initial setup screen appears. There, the user enters specific preferences, such as "I like traveling" or "I take lots of photos of my pet." The device converts this data into an appropriate format (e.g., JSON) and sends it to the server. The server analyzes the received data and stores it in a database.

[1039] Image generation

[1040] The server utilizes a generative AI model (e.g., DALL-E) to generate images based on the user's preferences. As a specific implementation example, a prompt such as "Based on the data that the user says 'I like traveling,' please generate photos of natural scenery at travel destinations" is input to the generative AI model. The server then assembles the generated images into a data packet and sends it to the user's device. The device then stores the received data in local storage and displays it in the album app.

[1041] User interaction and feedback

[1042] Users can browse the album using touch gestures and, if necessary, use voice input to request the generation of new images. For example, they could say, "Show me a new mountain climbing photo." The device receives the speech through its microphone and converts it into text using a speech recognition engine (e.g., Google Speech-to-Text). The text is sent to the server, which analyzes the request and generates a new prompt. For example, "Please generate a mountain climbing photo." The generated photo is sent back to the user's device, saved to local storage, and immediately displayed in the album app.

[1043] Continuous learning and optimization

[1044] The server collects user activity history, records which photos are most frequently viewed, and records which are most frequently requested, which is used to update the training dataset for the generative AI model and optimize the generation algorithm, so that subsequent image generation will be more in line with the user's preferences.

[1045] Specific examples

[1046] First-time setup

[1047] User: Launches the album app for the first time and enters preference information such as "I like traveling" and "I take lots of photos of my pets."

[1048] Terminal: Converts input data into JSON format and sends it to the server via an HTTP request.

[1049] Server: Analyzes the received data and stores it in a database. Prompts are input into the AI ​​model to generate images, such as travel photos and pet photos based on hobbies and preferences.

[1050] Device: The initially generated photos are saved to local storage and displayed in the album app.

[1051] Daily use

[1052] User: While browsing an album, make a voice request: "Show me new beach photos."

[1053] Device: Receives voice requests, converts them into text using a speech recognition engine, and sends the text to the server.

[1054] Server: Creates prompts to generate beach photos based on voice requests, inputs them into the generative AI model, and sends the generated beach photos to the user's device.

[1055] Device: Save the generated beach photos to local storage and display them instantly in the Album app.

[1056] The present invention allows users to easily obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" that are optimized to their individual preferences.

[1057] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1058] Step 1: Collect user data

[1059] User: Launches the album app for the first time and enters hobby and preference data (e.g., "I like traveling" or "I take lots of photos of my pets") on the initial setup screen.

[1060] Terminal: Converts the data entered by the user into JSON format and temporarily stores it in local storage.

[1061] Device: Sends the saved data to the server as an HTTP request.

[1062] Server: Analyzes the received data and stores it in the user database. For analysis, the Python pandas library is used to format the data and extract the necessary fields.

[1063] Input: User-entered information about your interests and preferences.

[1064] Output: Cleaned and formatted database entries.

[1065] Step 2: Image generation

[1066] Server: Based on the user's preference information, the server sends a prompt to the generative AI model (e.g., DALL-E). It creates a specific prompt such as, "Based on the data that the user says 'I like natural scenery,' please generate a photo of a natural scenery."

[1067] Server: Receives the image data returned by the generative AI model and converts it into an appropriate format, such as JPEG.

[1068] Server: Builds the image data into a data packet and sends it to the user's device as an HTTP response.

[1069] Input: User preference information, prompts for the generative AI model.

[1070] Output: The generated image data.

[1071] Step 3: Album View

[1072] Terminal: Analyzes the received image data and saves it to a local storage device in a common image format such as JPEG.

[1073] On the device: Generate thumbnails of the images and display them in the album app. For example, use the Pillow library to generate thumbnails.

[1074] Input: Image data sent from the server.

[1075] Output: Image files saved to local storage, thumbnail display in app.

[1076] Step 4: User interaction and feedback

[1077] 4.1 Album browsing

[1078] User: Browse albums within the app using touch controls.

[1079] Device: Detects touch input and enlarges the image according to the touch position. The image enlargement process uses the built-in graphics library.

[1080] Input: Touch input data.

[1081] Output: The enlarged image.

[1082] 4.2 Voice Requests

[1083] User: Requests the generation of a new image via speech input (e.g., "Show me a new mountain climbing photo").

[1084] Device: Voice input is received via a microphone and converted into text using a speech recognition engine (e.g., Google Speech-to-Text).

[1085] Terminal: The converted text data is sent to the server as an HTTP request.

[1086] Input: The user's voice request.

[1087] Output: Data converted to text by the speech recognition engine and sent as an HTTP request.

[1088] 4.3 Creating and displaying new images

[1089] Server: Analyzes the voice request and inputs a new prompt to the generative AI model. For example, send a prompt such as "Generate a photo of mountain climbing."

[1090] Server: Sends the generated new image data to the user's terminal.

[1091] Device: New images received are saved to local storage and immediately displayed in the Album app.

[1092] Input: A prompt to the generative AI model via a voice request.

[1093] Output: New image data generated, new image displayed.

[1094] Step 5: Continue learning and optimization

[1095] Server: Collects user activity history, including browsing frequency, type of image requested, etc.

[1096] Server: Updates the training dataset for the generative AI model based on the collected operation history and optimizes the image generation algorithm. For updates, machine learning libraries such as Scikit-learn and TensorFlow are used.

[1097] Input: User operation history.

[1098] Output: An optimized generative AI model.

[1099] Through these steps, users can easily obtain and enjoy high-quality "happy false memories" based on their preferences.

[1100] (Application example 1)

[1101] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1102] Conventional photo album applications create albums based on photos that users have actually taken, so unless users have experience taking photos, the albums are not complete. Also, it is difficult for users to easily collect photos that suit their preferences, which limits the user experience.

[1103] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1104] In this invention, the server includes means for generating images using a generative AI model based on the user's hobby and preference data, means for sending the generated images to the user's device, and means for updating the generation algorithm using prompts based on the user's operation history. This allows the user to easily enjoy "happy false memories" that were not captured, and a variety of photos and videos that suit the user's preferences can be generated and provided in real time.

[1105] "User's hobby and preference data" is information entered by the user that represents the user's interests, concerns, favorite activities, and subjects.

[1106] "Server" means a computer system that receives, processes, stores, and generates data submitted by users.

[1107] A "generative AI model" is an algorithm that uses artificial intelligence to generate new images and videos based on user preferences and taste data.

[1108] The "means for generating images" refers to a mechanism that uses a generative AI model to create images based on the user's taste and preference data.

[1109] The "means for transmitting images to the user's terminal" is a communication protocol for transferring the generated image data to the device used by the user.

[1110] "Means for displaying images" refers to the functionality for visually displaying images generated on the user's device.

[1111] "User operation history" refers to records of touch operations, voice input, browsing history, etc., when a user uses an application.

[1112] A "prompt" is a textual instruction that instructs a generative AI model to generate a specific image or video.

[1113] "Speech recognition technology" is a technology that interprets and processes a user's verbal instructions by analyzing voice data and converting it into text data.

[1114] "Local storage" is a data storage area built into the user's device.

[1115] The present invention relates to a content distribution system for providing "happy false memories" that are not captured by the user, and is implemented as follows.

[1116] Device and Software Configuration

[1117] The hardware used is smartphones, smart glasses, and head-mounted displays (HMDs), such as Google Glass, North Focals, Oculus Rift, and HTC Vive. The software used is Google Cloud Speech-to-Text for voice recognition, generative AI models (e.g., OpenAI's DALL-E, Stable Diffusion) for image generation, and Firebase for the database.

[1118] User Data Collection

[1119] When users first launch the application, they enter their interests and preferences, which can include specific details like "I like to travel" or "I take lots of photos of my pets." This data is sent from their smartphone or device to the cloud and stored in a Firebase database.

[1120] Generate initial album

[1121] The server retrieves the user's interest data from the Firebase database and passes prompts to the generative AI model to generate an image. For example, if a user enters "I like natural scenery," the generative AI model receives the prompt "Generate a photo of a peaceful mountain hike with clear skies and lush greenery." The generated image is constructed as a data packet and sent to the user's device. The received data is stored in local storage on the device, and the image is displayed.

[1122] User Actions

[1123] Users can request the generation of new images using touch or voice input. For example, by saying, "Show me a new beach photo," the voice data is converted to text through Google Cloud Speech-to-Text and sent to the server. The server then provides the generative AI model with a new prompt, which generates a new photo. This process instantly sends the newly generated image to the user's device and displays it.

[1124] Continuous learning and optimization

[1125] The server collects user operation history, records which photos are viewed most frequently, and records which are frequently requested. Based on the collected operation history, the algorithm of the generative AI model is updated and optimized, so that future image generation will be more in line with the user's preferences.

[1126] Specific examples

[1127] For example, if a user makes a voice request such as "Show me new mountain climbing photos," the following prompt sentence is passed to the generative AI model:

[1128] "Generate a photo of a peaceful mountain hike with clear skies and lush greenery"

[1129] Through the above process, users can create and enjoy a variety of photos and videos that suit their tastes in real time.

[1130] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1131] Step 1:

[1132] When a user launches the application for the first time, they enter their hobbies and preferences, such as "I like traveling" or "I take lots of photos of my pets." The entered data is sent from the smartphone or device to the server. The data is then temporarily processed on the user's device and converted into the appropriate format.

[1133] Step 2:

[1134] The server receives the interest and preference data sent by the user and stores it in the Firebase database. The input here is the interest and preference data, and the output is the stored data. The server analyzes the received data and prepares the information necessary for later image generation.

[1135] Step 3:

[1136] The server retrieves interest and preference data from the Firebase database and generates an image by passing a prompt to the generative AI model. For example, a specific prompt such as "Generate a photo of a peaceful mountain hike with clear skies and lush greenery" is provided to the generative AI model. The input is the interest and preference data and the prompt, and the output is the generated image.

[1137] Step 4:

[1138] The generated image is sent as a data packet from the server to the user's device. The input here is the generated image data, and the output is the image data received by the user's terminal. The server converts the image data into an appropriate format and transmits it via a communication protocol.

[1139] Step 5:

[1140] The device stores the image data received from the server in local storage and displays the image, where the input is the received image data and the output is the image visually displayed to the user. The device displays the stored image data in a user interface in an appropriate format.

[1141] Step 6:

[1142] The user requests the generation of a new image using touch or voice input. For example, they might say, "Show me a new beach photo." This voice data is picked up by the device's microphone and converted to text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data.

[1143] Step 7:

[1144] The device sends a voice request, converted into text, to the server, which analyzes the request and generates a new prompt. For example, a prompt such as "Generate a photo of a sunny beach with clear blue water" is passed to a generative AI model. The input is text data, and the output is the new prompt.

[1145] Step 8:

[1146] The server provides the new prompt to the generative AI model and generates a new image. The generated image is then sent back to the user's device. The input is the new prompt and the generative AI model, and the output is the newly generated image.

[1147] Step 9:

[1148] The device saves the newly received image to local storage and immediately displays it to the user. Here, the input is the newly received image data and the output is the image displayed to the user, allowing the user to immediately view the new content they requested.

[1149] Step 10:

[1150] The server continuously collects user operation history, recording which photos are viewed most frequently and which requests are most frequently made. This data is saved as operation history. The input is the user operation data, and the output is the saved operation history.

[1151] Step 11:

[1152] The server updates and optimizes the generative AI model algorithm based on the collected operation history. For example, if there are many requests for a particular situation, the server strengthens the generative algorithm to accommodate that situation. The input is operation history data, and the output is an optimized generative AI model. This process improves image generation from the next time onwards to better match the user's preferences.

[1153] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1154] The present invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that are not actually taken, and is implemented as follows.

[1155] User Data Collection

[1156] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1157] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[1158] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1159] Introducing the Emotion Engine

[1160] Device: The device acquires the user's facial expressions and tone of voice, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server along with the user data.

[1161] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[1162] Generate initial album

[1163] Server: Generates images using AI technology based on the user's hobby, preference, and emotional data. For example, if a user is recognized as "liking natural scenery" and "relaxed," a photo of a serene natural landscape will be generated.

[1164] Server: Assembles the generated image as a data packet and sends it to the user's device.

[1165] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[1166] User interaction and feedback

[1167] Users can browse albums with touch or use voice input to request new images (e.g., "Show me pictures of the beach").

[1168] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[1169] Server: Analyzes the voice request and extracts parameters for generating new images. For example, based on the instruction "Photo of mountain climbing," it generates an appropriate photo and incorporates emotional data.

[1170] Server: Sends the generated new image to the user device.

[1171] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1172] Continuous learning and optimization

[1173] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[1174] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[1175] Specific examples

[1176] First-time setup

[1177] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take a lot of photos of my pets."

[1178] Terminal: Sends input data and the user's initial emotion data to the server.

[1179] Server: Analyzes hobby and preference data and emotional data to generate appropriate travel and pet photos.

[1180] Terminal: Displays the initially generated photo.

[1181] Daily use

[1182] User: While browsing an album, say "Show me new beach photos."

[1183] Device: Sends voice requests and emotion data to the server.

[1184] Server: Generates and sends beach photos based on the user's request and emotion data.

[1185] Device: Generated beach photos are saved to local storage and displayed immediately.

[1186] This system allows users to easily and intuitively experience "happy false memories" that match a specific emotional state.

[1187] The processing flow will be explained below.

[1188] Step 1:

[1189] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1190] Step 2:

[1191] Terminal: Temporarily stores the data entered by the user and sends it to the server, where it is converted into the appropriate format.

[1192] Step 3:

[1193] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1194] Step 4:

[1195] Device: Captures the user's facial expressions and voice tone, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server.

[1196] Step 5:

[1197] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[1198] Step 6:

[1199] Server: Generates images for the initial album using AI technology based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," it generates photos of tranquil natural scenery.

[1200] Step 7:

[1201] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[1202] Step 8:

[1203] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[1204] Step 9:

[1205] User: Browse the album with touch input or use voice input to request the generation of new images, for example, "Show me new beach photos."

[1206] Step 10:

[1207] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[1208] Step 11:

[1209] Server: Receives voice requests and emotion data, analyzes them, and extracts new image generation parameters. For example, in response to a request for a "beach photo," it generates an appropriate beach photo, taking into account the user's emotion data.

[1210] Step 12:

[1211] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[1212] Step 13:

[1213] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1214] Step 14:

[1215] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[1216] Step 15:

[1217] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[1218] By following the above steps, users can easily and intuitively experience a "happy false memory" that matches their specific emotional state.

[1219] Example 2

[1220] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1221] With the recent advancement of digital technology, users are increasingly seeking tools that allow them to freely enjoy their imaginations and memories. However, existing album applications only collect and manage memories that users have clearly experienced, lacking the ability to generate new memories and experiences. Furthermore, they lack the ability to generate and display content that takes into account the user's emotional state, making it difficult to provide an experience that meets the user's individual needs and emotions. Furthermore, with few systems offering intuitive operation through voice input or real-time feedback, users are often forced to use complex operations.

[1222] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting user interest and preference data, means for transmitting user data to the server, means for acquiring user emotion data, means for generating an image based on the received user data and emotion data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting user operation history, means for updating the generation algorithm based on the operation history and emotion data, means for interpreting the user's request using voice recognition and generating a new image, and means for saving the generated image in local storage and immediately displaying it to the user. This makes it possible to easily generate new memories and experiences according to the user's interest and preference and emotional state, and intuitively provide content suitable for each individual user.

[1223] "User" means an individual or legal entity that uses the System.

[1224] "Hobby and preference data" refers to information about the genres and preferences that a user is interested in.

[1225] "Emotional data" refers to information about a user's emotional state obtained from facial expressions, tone of voice, etc.

[1226] "Server" refers to a computer system that processes, stores, and analyzes data submitted by users.

[1227] "Terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user.

[1228] "Received User Data" refers to information about a User that the Server receives from a Terminal.

[1229] "Image generation" refers to the process of creating new images using AI technology.

[1230] "Display means" refers to the method or technology that allows the generated image to be visually viewed on the user's device.

[1231] "Operation history" refers to the record of various operations performed by a user using the system.

[1232] "Generation algorithm" refers to the calculation procedure or model for generating an image based on received data.

[1233] A "voice recognition engine" refers to software that analyzes a user's voice and converts it into text.

[1234] "Local storage" refers to the data storage area within the user's device.

[1235] This invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that they have not taken. This system provides users with newly generated images using AI technology based on the user's taste data and emotion data.

[1236] Hardware and software used

[1237] Device: A device that is directly operated by the user, such as a smartphone, tablet, or personal computer, is used. The device is equipped with a camera and microphone, which are used to capture the user's facial expressions and voice.

[1238] Server: Use a high-performance cloud server or data center. The server performs central processing such as data storage, analysis, and image generation. Specific examples include servers provided by major cloud service providers (e.g., Amazon Web Services, Google Cloud Platform).

[1239] Software: The main software components of this system include an emotion engine, a speech recognition engine, and a generative AI model. We use OpenCV for the emotion engine and Google Cloud Speech-to-Text for the speech recognition engine. We use generative AI models (e.g., DALL·E, Stable Diffusion) for image generation.

[1240] Data flow and processing

[1241] 1. User data input

[1242] When a user launches the album app for the first time, they enter their name and hobbies (e.g., "travel," "nature scenery," "pet photos," etc.). The device temporarily stores this data, converts it into an appropriate format, and sends it to the server.

[1243] 2. Analysis and saving of initial setting data

[1244] The server analyzes the received user data and stores it in a database. During this analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[1245] 3. Acquiring and sending emotion data

[1246] The device's camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This emotion data is sent to a server, where it is analyzed and saved.

[1247] 4. Initial album generation

[1248] The server generates images using a generative AI model based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," a photo of a calm natural scenery will be generated. An example of a prompt sentence is "relaxing natural scenery."

[1249] 5. Sending and displaying images

[1250] The server then assembles the generated images into a data packet and sends it to the user's device, which receives it, stores it in local storage, and displays it in the album app.

[1251] User interaction and feedback

[1252] Users can browse the album with touch gestures and use voice input to request the generation of new images (e.g., "Show me new beach photos.") The device receives the voice request, converts it into text using a speech recognition engine, and sends it to the server.

[1253] The server analyzes the voice request and extracts parameters for generating a new image. For example, a prompt such as "beach photo" is input into the generative AI model, which then generates an image incorporating the user's emotional data. The image is then sent to the device, stored in local storage, and instantly displayed to the user.

[1254] Continuous learning and optimization

[1255] The server continuously collects the user's operation history and emotional data, and updates and optimizes the generation algorithm based on this data, so that the next image generation will be more in line with the user's tastes, preferences, and emotional state.

[1256] As described above, the present invention is a system that provides "happy false memories" that take into account the individual preferences and emotions of the user, and can provide the user with a new experience in an intuitive and simple manner.

[1257] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1258] Step 1: User Data Input

[1259] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1260] Input: User name, hobbies and preferences

[1261] Output: Formatted user data (e.g., JSON format)

[1262] Specific behavior: The user enters their name and hobbies and preferences into the text boxes and check boxes, then presses the submit button.

[1263] Step 2: Send data

[1264] Terminal: Temporarily stores the input data, converts it into an appropriate format (e.g., JSON format), and sends it to the server.

[1265] Input: User data (name, hobbies, preferences)

[1266] Output: User data sent to the server

[1267] Specific behavior: After format conversion, the data is sent to the server using an HTTP request.

[1268] Step 3: Analyze and save the initial setup data

[1269] Server: Analyzes the received user data and stores it in a database. During the analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[1270] Input: Received user data

[1271] Output: Parsed user data stored in a database

[1272] Specific operation: Apply data analysis algorithms to generate parameters based on the user's preferences. Store the analysis results in a database.

[1273] Step 4: Obtaining and sending emotion data

[1274] Device: The camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This data is sent to the server.

[1275] Input: User's facial expression data, voice tone data

[1276] Output: Emotion data sent to the server

[1277] Specific operation: The camera captures faces in real time and records voice input. These data are analyzed by the emotion engine, and emotion data is extracted and sent to the server.

[1278] Step 5: Generate the initial album

[1279] Server: Generates images using a generative AI model based on the user's taste and emotion data.

[1280] Input: Interest data, emotion data, prompt sentence (e.g., "Relaxing nature scenery")

[1281] Output: Generated image data

[1282] Specific operation: A prompt sentence and emotion data are input to the generative AI model, and the AI ​​generates an image. The generated image is converted into a data packet.

[1283] Step 6: Send and view images

[1284] Server: Assembles the generated image as a data packet and sends it to the user's device.

[1285] Device: The received image data is stored in local storage and displayed to the user in the album app.

[1286] Input: Generated image data

[1287] Output: Image displayed in the album app

[1288] Specific behavior: Receive image data from the server, save it to local storage, and display the image in the app's user interface.

[1289] Step 7: Browse albums and request new images

[1290] User: Browse albums with touch and use voice input to request new images (e.g., "Show me new beach photos").

[1291] Input: Touch, voice requests

[1292] Output: New image request data

[1293] Specific actions: Swipe to change images or use the microphone to input voice commands.

[1294] Step 8: Sending and analyzing audio data

[1295] Terminal: Receives voice requests, converts them into text using a speech recognition engine, and sends them to the server.

[1296] Input: Voice request

[1297] Output: The request text sent to the server

[1298] Specific operation: Converts speech to text and sends the text data to the server.

[1299] Step 9: Generate and send a new image

[1300] Server: Analyzes the voice request, extracts parameters for generating a new image, generates the image using a generative AI model, and sends it to the user's device.

[1301] Input: Voice request text, emotion data

[1302] Output: The new image data generated.

[1303] Specific operation: Analyzes the voice request and generates a prompt (e.g., "New beach photo"). Generates an image using a generative AI model, converts it into a data packet, and sends it to the device.

[1304] Step 10: Displaying the new image

[1305] On the device: The new image is generated and saved to local storage, and is immediately displayed to the user.

[1306] Input: Newly generated image data

[1307] Output: The new image displayed in the Album app

[1308] Specific behavior: Receives image data from the server, saves it to local storage, and displays the image in the album app's user interface.

[1309] Step 11: Continue learning and optimization

[1310] Server: Continuously collects user operation history and emotion data, and updates and optimizes the generation algorithm.

[1311] Input: Operation history, emotion data

[1312] Output: Updated and optimized generation algorithm

[1313] Specific operation: Obtain operation history and emotion data from the database. Based on this data, the generation algorithm is retrained and optimized.

[1314] (Application example 2)

[1315] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1316] While existing virtual stores recommend products based on users' preferences, they do not use real-time emotional data to recommend products that reflect the user's current emotional state. As a result, the accuracy of product recommendations to users is low, limiting the improvement of the user experience. In addition, systems linked to voice recognition have not been fully utilized when users make specific requests, making them difficult to use intuitively.

[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1318] In this invention, the server includes means for inputting user interest and preference data, means for transmitting the user data to the server, means for generating an image based on the received data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for acquiring and analyzing the user's emotional data in real time, means for recommending products based on the acquired emotional data, and means for displaying the products on the terminal. This enables personalized product recommendations based on the user's emotional state in real time, improving the user experience and enabling intuitive operation.

[1319] "Hobbies and Preference Data" is information that indicates a user's interests, favorite activities, and hobbies.

[1320] "Emotional data" is data that represents the emotional state of a user, obtained from the user's facial expression, tone of voice, etc.

[1321] An "image generation means" is an algorithm or software that creates an image based on the data received.

[1322] The "product recommendation means" is a system that suggests products suitable for the user based on emotional data and hobby and preference data.

[1323] A "server" is a remote computer or system that has functions such as receiving, analyzing, and storing data, generating images, and recommending products.

[1324] "Terminal" means a device that can be directly operated by a user, including a smartphone, smart glasses, tablet, etc.

[1325] "Operation history" refers to historical data of operations and actions taken by a user when using the system.

[1326] "Speech recognition" is a technology that converts a user's voice into text data and understands the instructions based on that.

[1327] "Real-time" refers to data acquisition, analysis, and response occurring immediately and without delay.

[1328] This invention is a virtual store system that allows users to acquire their own emotional data in real time using their smart devices and recommends personalized products based on that data. To implement this invention, the following means are required.

[1329] First, the user puts on the smart glasses and starts the application. At startup, the user enters their name and hobby / preference data (e.g., "outdoor goods" or "casual fashion"). This data is temporarily stored in the smart glasses and sent to the server.

[1330] Next, the server prepares parameters for generating the initial album based on the received user's tastes and preferences. It also acquires the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine. The analyzed emotion data is sent to the server and stored in a database.

[1331] The server combines the user's taste and preference data with real-time emotional data to generate an image using a generative AI model. For example, if the user is in a "relaxed" state and prefers "outdoor goods," an image of outdoor goods with a relaxed atmosphere will be generated. The generated image and product information are organized into a data packet and sent to the user's smart glasses.

[1332] The smart glasses display the received image data and product information. The user can use voice requests to instruct the server to generate additional images. The voice requests are converted into text by a speech recognition engine, and the data is sent to the server. The server analyzes the voice requests and makes appropriate product recommendations.

[1333] The software components of this application include:

[1334] OpenCV: Used to analyze user facial expressions.

[1335] EmotionEngine: A custom AI model that identifies user emotions in real time.

[1336] ProductRecommender: An algorithm that provides personalized product recommendations based on user sentiment data.

[1337] ARDisplay: An augmented reality library that displays products virtually in the field of view of smart glasses.

[1338] As a concrete example, when a user enters a shopping mall and shows a relaxed expression, the emotion engine recognizes the state as "relaxed." Based on this, the server recommends "casual fashion items" and displays them on the smart glasses' display. When the user makes a voice request such as "Show me a new casual shirt," the voice is converted into text data and sent to the server. Based on this request, the server generates an image of an appropriate shirt and sends it back to the smart glasses.

[1339] Example prompt sentence:

[1340] Prompt: Recommend products suitable for the user when they are relaxed. The user is a man in his 30s who likes outdoor gear.

[1341] Example output: Recommendations include "Casual T-shirt", "Outdoor backpack", and "Sports sunglasses".

[1342] This enables personalized product suggestions based on user sentiment in real time, improving the user experience and providing an intuitive and convenient shopping experience.

[1343] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1344] Step 1:

[1345] Input of user's hobby and preference data

[1346] How it works: The user puts on the smart glasses and launches the application. On first launch, the user enters their name and preferences (e.g., "outdoor gear" or "casual fashion").

[1347] Input: User's basic information and hobbies and preferences

[1348] Data processing: The input data is temporarily stored in the smart glasses and converted into the required format.

[1349] Output: Formatted user data

[1350] Step 2:

[1351] Sending user data

[1352] How it works: The smart glasses send formatted user data to the server.

[1353] Input: Formatted user data

[1354] Data processing: None

[1355] Output: User data sent to the server

[1356] Step 3:

[1357] First, prepare to generate an album based on hobby and preference data.

[1358] Operation: The server analyzes the received user data and prepares parameters for generating the initial album based on the user's tastes and preferences.

[1359] Input: User data sent to the server

[1360] Data Calculation: Analyzing user data

[1361] Output: Parameters for initial album generation

[1362] Step 4:

[1363] Emotion data collection and analysis

[1364] How it works: The smart glasses capture the user's facial expressions and voice tone in real time, analyze these data with the emotion engine, and send the emotion data to the server.

[1365] Input: User facial expressions and voice tone

[1366] Data Computation: Emotion Analysis with Emotion Engine

[1367] Output: Emotion data

[1368] Step 5:

[1369] Emotion data storage and analysis

[1370] How it works: The server analyzes the received emotion data and stores it in a database.

[1371] Input: Emotion data

[1372] Data Computing: Sentiment Data Analysis and Storage

[1373] Output: Emotion data stored in a database

[1374] Step 6:

[1375] Image and product recommendation generation

[1376] How it works: By combining user preference data with real-time emotional data, a generative AI model is used to generate images and prepare product information.

[1377] Input: Parameters for initial album generation, emotion data

[1378] Data processing: Image generation using generative AI models, creation of product information

[1379] Output: Generated image data and product information

[1380] Step 7:

[1381] Submitting images and product information

[1382] Operation: The server assembles the generated image data and product information into a data packet and sends it to the user's smart glasses.

[1383] Input: Generated image data and product information

[1384] Data processing: Data packet construction

[1385] Output: Image data and product information sent to smart glasses

[1386] Step 8:

[1387] Display of images and product information

[1388] Operation: The smart glasses display the received image data and product information.

[1389] Input: Image data and product information sent to smart glasses

[1390] Data processing: None

[1391] Output: Image and product information displayed to the user

[1392] Step 9:

[1393] Processing user requests

[1394] How it works: A user uses a voice request to request additional images or product information. The smart glasses convert the voice request into text and send it to the server.

[1395] Input: User's voice request

[1396] Data calculation: Text conversion using a speech recognition engine

[1397] Output: Text request sent to the server

[1398] Step 10:

[1399] Generate and submit new images and products

[1400] How it works: The server analyzes the voice request, generates the appropriate images and products, and sends them to the smart glasses.

[1401] Input: The text request sent to the server

[1402] Data processing: Generative AI models generate new images and suggest products, and create data packets

[1403] Output: New image and product information sent to the smart glasses

[1404] Step 11:

[1405] New images and product displays

[1406] How it works: The smart glasses display the new images and product information they receive.

[1407] Input: New image and product information sent to the smart glasses

[1408] Data processing: None

[1409] Output: The new image and product information displayed to the user

[1410] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1412] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1413] [Fourth embodiment]

[1414] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1415] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1416] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1417] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1418] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1420] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1421] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1422] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1423] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1424] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1425] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1426] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1427] The present invention relates to an album application for providing "fake happy memories" that are not taken by the user, and is implemented as follows.

[1428] User Data Collection

[1429] User: When launching the album app for the first time, the user enters their hobbies and interests on the initial setup screen, including specific information such as "I like traveling" and "I take lots of photos of my pets."

[1430] Terminal: Processes the data entered by the user and sends it to the server, where it is converted into an appropriate format and temporarily stored.

[1431] Server: Analyzes the received user data and stores it in a database, which provides the necessary information for later image generation.

[1432] Generate initial album

[1433] Server: Generates images using AI technology based on the user's hobbies and preferences. For example, if a user enters data such as "I like natural scenery," the AI ​​will generate an appropriate photo of a natural scenery.

[1434] Server: Assembles the generated image as a data packet and sends it to the user's device.

[1435] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[1436] User interaction and feedback

[1437] User: Browse albums with touch or use voice input to request new images, for example, "Show me my new mountain climbing photos."

[1438] On your device: Detects touch input and displays details of selected photos, and sends voice input to a speech recognition engine for conversion to text.

[1439] Server: Receives and analyzes voice requests to extract new image generation requirements. For example, following the instruction "mountain climbing photos," it generates new mountain climbing photos.

[1440] Server: Sends the generated new image to the user device.

[1441] On the device: New images are received and saved to local storage, and are immediately displayed to the user.

[1442] Continuous learning and optimization

[1443] Server: Collects user activity history and records which photos are most frequently viewed and which requests are most frequently made.

[1444] Server: Using the collected operation history, the generation algorithm is updated and optimized, so that subsequent image generation will better match the user's preferences.

[1445] Specific examples

[1446] First-time setup

[1447] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take lots of photos of my pets."

[1448] Terminal: Sends input data to the server.

[1449] Server: Analyzes the received data and generates travel photos and pet photos.

[1450] Terminal: Displays the initially generated photo.

[1451] Daily use

[1452] User: While browsing an album, say "Show me new beach photos."

[1453] Device: Sends voice requests to the server.

[1454] Server: Generates beach photos and sends them to the user device.

[1455] Device: Generated beach photos are saved to local storage and displayed immediately.

[1456] In this way, users can enjoy the experience of gaining new, happy "false memories." By using this system, users can easily enjoy a variety of photos that suit their preferences.

[1457] The processing flow will be explained below.

[1458] Step 1:

[1459] User: Launch the Album app. On the initial setup screen, enter your name and hobbies (e.g., "Travel," "Nature Scenery," "Pet Photos").

[1460] Step 2:

[1461] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[1462] Step 3:

[1463] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1464] Step 4:

[1465] Server: Based on the prepared parameters, images for the initial album are generated using a generative AI (e.g., GAN or VQ-VAE-2). The generated images are saved as temporary files.

[1466] Step 5:

[1467] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[1468] Step 6:

[1469] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[1470] Step 7:

[1471] User: Browse albums with touch input, select photos to view and navigate to details, and use voice input to request new photos (e.g., "Show me photos of the beach").

[1472] Step 8:

[1473] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, uses a speech recognition engine to convert speech into text and send it to the server.

[1474] Step 9:

[1475] Server: Analyzes the voice request and extracts parameters for generating a new image. Based on the extracted parameters, a new image is generated using generative AI.

[1476] Step 10:

[1477] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[1478] Step 11:

[1479] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1480] Step 12:

[1481] Server: Continuously collects user activity history, recording which photos are most frequently viewed and what requests are most frequently made.

[1482] Step 13:

[1483] Server: Updates and optimizes the generation algorithm based on the collected operation history, so that subsequent image generation can be more tailored to the user's preferences.

[1484] By following these steps, users can easily and intuitively operate the app to create new "happy false memories."

[1485] Example 1

[1486] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1487] In modern digital album applications, many users spend time and effort efficiently collecting photos that match their hobbies and preferences. However, there are limited ways for users to easily obtain high-quality images that meet their individual needs without having to go through the trouble of manually collecting and selecting content. Furthermore, there are no systems that automatically optimize generated images using user operation history. Furthermore, the ability to interpret user requests using voice recognition and respond immediately is also insufficient. Therefore, the objective of this invention is to provide a system that provides "happy false memories" that match the user's preferences and are continuously optimized.

[1488] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1489] In this invention, the server includes means for inputting user preference information, means for transmitting the preference information to the server, means for generating images based on the received preference information, means for transmitting the generated images to the user's device, means for displaying the images transmitted to the device, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for interpreting the user's request using voice recognition and generating images, and means for saving the generated images in a local storage device and immediately displaying them to the user. This allows the user to easily and efficiently obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" optimized for their individual preferences.

[1490] "User" refers to an individual who uses the Album Application.

[1491] "Preference information" refers to data related to a user's hobbies and preferences that is entered on the initial setup screen, etc.

[1492] A "server" refers to a computing system that receives data from users and performs a series of processes such as analysis, image generation, and database management.

[1493] "Device" refers to the end-user devices used by a User, such as a smartphone, tablet, or PC.

[1494] "Means for generating images" refers to methods or functions for creating new images using a generative AI model based on received preference information.

[1495] "Voice recognition" refers to the technology that converts requests input by voice by the user into text data.

[1496] "Operation history" refers to data and logs generated by operations performed when a user uses the album application.

[1497] "Means for updating the generation algorithm" refers to a method for updating the training data of the AI ​​model based on collected operation history to improve the accuracy and efficiency of the image generation algorithm.

[1498] "Local storage device" refers to the data storage area that exists within the user's device, and specifically includes internal storage and external memory.

[1499] "Happy false memories" refer to fictional photos or image data that bring about a sense of happiness based on the user's preferences or requests, even though the user has not actually experienced them.

[1500] This invention relates to an album application that generates and provides "happy false memories" based on user preference information. This system involves a series of processes that collect user preference information, generate images using a generative AI model based on that information, and continuously optimize the images based on the user's operation history.

[1501] User Data Collection

[1502] When a user launches the album application for the first time, an initial setup screen appears. There, the user enters specific preferences, such as "I like traveling" or "I take lots of photos of my pet." The device converts this data into an appropriate format (e.g., JSON) and sends it to the server. The server analyzes the received data and stores it in a database.

[1503] Image generation

[1504] The server utilizes a generative AI model (e.g., DALL-E) to generate images based on the user's preferences. As a specific implementation example, a prompt such as "Based on the data that the user says 'I like traveling,' please generate photos of natural scenery at travel destinations" is input to the generative AI model. The server then assembles the generated images into a data packet and sends it to the user's device. The device then stores the received data in local storage and displays it in the album app.

[1505] User interaction and feedback

[1506] Users can browse the album using touch gestures and, if necessary, use voice input to request the generation of new images. For example, they could say, "Show me a new mountain climbing photo." The device receives the speech through its microphone and converts it into text using a speech recognition engine (e.g., Google Speech-to-Text). The text is sent to the server, which analyzes the request and generates a new prompt. For example, "Please generate a mountain climbing photo." The generated photo is sent back to the user's device, saved to local storage, and immediately displayed in the album app.

[1507] Continuous learning and optimization

[1508] The server collects user activity history, records which photos are most frequently viewed, and records which are most frequently requested, which is used to update the training dataset for the generative AI model and optimize the generation algorithm, so that subsequent image generation will be more in line with the user's preferences.

[1509] Specific examples

[1510] First-time setup

[1511] User: Launches the album app for the first time and enters preference information such as "I like traveling" and "I take lots of photos of my pets."

[1512] Terminal: Converts input data into JSON format and sends it to the server via an HTTP request.

[1513] Server: Analyzes the received data and stores it in a database. Prompts are input into the AI ​​model to generate images, such as travel photos and pet photos based on hobbies and preferences.

[1514] Device: The initially generated photos are saved to local storage and displayed in the album app.

[1515] Daily use

[1516] User: While browsing an album, make a voice request: "Show me new beach photos."

[1517] Device: Receives voice requests, converts them into text using a speech recognition engine, and sends the text to the server.

[1518] Server: Creates prompts to generate beach photos based on voice requests, inputs them into the generative AI model, and sends the generated beach photos to the user's device.

[1519] Device: Save the generated beach photos to local storage and display them instantly in the Album app.

[1520] The present invention allows users to easily obtain high-quality images based on their preferences, enabling them to continuously experience "happy false memories" that are optimized to their individual preferences.

[1521] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1522] Step 1: Collect user data

[1523] User: Launches the album app for the first time and enters hobby and preference data (e.g., "I like traveling" or "I take lots of photos of my pets") on the initial setup screen.

[1524] Terminal: Converts the data entered by the user into JSON format and temporarily stores it in local storage.

[1525] Device: Sends the saved data to the server as an HTTP request.

[1526] Server: Analyzes the received data and stores it in the user database. For analysis, the Python pandas library is used to format the data and extract the necessary fields.

[1527] Input: User-entered information about your interests and preferences.

[1528] Output: Cleaned and formatted database entries.

[1529] Step 2: Image generation

[1530] Server: Based on the user's preference information, the server sends a prompt to the generative AI model (e.g., DALL-E). It creates a specific prompt such as, "Based on the data that the user says 'I like natural scenery,' please generate a photo of a natural scenery."

[1531] Server: Receives the image data returned by the generative AI model and converts it into an appropriate format, such as JPEG.

[1532] Server: Builds the image data into a data packet and sends it to the user's device as an HTTP response.

[1533] Input: User preference information, prompts for the generative AI model.

[1534] Output: The generated image data.

[1535] Step 3: Album View

[1536] Terminal: Analyzes the received image data and saves it to a local storage device in a common image format such as JPEG.

[1537] On the device: Generate thumbnails of the images and display them in the album app. For example, use the Pillow library to generate thumbnails.

[1538] Input: Image data sent from the server.

[1539] Output: Image files saved to local storage, thumbnail display in app.

[1540] Step 4: User interaction and feedback

[1541] 4.1 Album browsing

[1542] User: Browse albums within the app using touch controls.

[1543] Device: Detects touch input and enlarges the image according to the touch position. The image enlargement process uses the built-in graphics library.

[1544] Input: Touch input data.

[1545] Output: The enlarged image.

[1546] 4.2 Voice Requests

[1547] User: Requests the generation of a new image via speech input (e.g., "Show me a new mountain climbing photo").

[1548] Device: Voice input is received via a microphone and converted into text using a speech recognition engine (e.g., Google Speech-to-Text).

[1549] Terminal: The converted text data is sent to the server as an HTTP request.

[1550] Input: The user's voice request.

[1551] Output: Data converted to text by the speech recognition engine and sent as an HTTP request.

[1552] 4.3 Creating and displaying new images

[1553] Server: Analyzes the voice request and inputs a new prompt to the generative AI model. For example, send a prompt such as "Generate a photo of mountain climbing."

[1554] Server: Sends the generated new image data to the user's terminal.

[1555] Device: New images received are saved to local storage and immediately displayed in the Album app.

[1556] Input: A prompt to the generative AI model via a voice request.

[1557] Output: New image data generated, new image displayed.

[1558] Step 5: Continue learning and optimization

[1559] Server: Collects user activity history, including browsing frequency, type of image requested, etc.

[1560] Server: Updates the training dataset for the generative AI model based on the collected operation history and optimizes the image generation algorithm. For updates, machine learning libraries such as Scikit-learn and TensorFlow are used.

[1561] Input: User operation history.

[1562] Output: An optimized generative AI model.

[1563] Through these steps, users can easily obtain and enjoy high-quality "happy false memories" based on their preferences.

[1564] (Application example 1)

[1565] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1566] Conventional photo album applications create albums based on photos that users have actually taken, so unless users have experience taking photos, the albums are not complete. Also, it is difficult for users to easily collect photos that suit their preferences, which limits the user experience.

[1567] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1568] In this invention, the server includes means for generating images using a generative AI model based on the user's hobby and preference data, means for sending the generated images to the user's device, and means for updating the generation algorithm using prompts based on the user's operation history. This allows the user to easily enjoy "happy false memories" that were not captured, and a variety of photos and videos that suit the user's preferences can be generated and provided in real time.

[1569] "User's hobby and preference data" is information entered by the user that represents the user's interests, concerns, favorite activities, and subjects.

[1570] "Server" means a computer system that receives, processes, stores, and generates data submitted by users.

[1571] A "generative AI model" is an algorithm that uses artificial intelligence to generate new images and videos based on user preferences and taste data.

[1572] The "means for generating images" refers to a mechanism that uses a generative AI model to create images based on the user's taste and preference data.

[1573] The "means for transmitting images to the user's terminal" is a communication protocol for transferring the generated image data to the device used by the user.

[1574] "Means for displaying images" refers to the functionality for visually displaying images generated on the user's device.

[1575] "User operation history" refers to records of touch operations, voice input, browsing history, etc., when a user uses an application.

[1576] A "prompt" is a textual instruction that instructs a generative AI model to generate a specific image or video.

[1577] "Speech recognition technology" is a technology that interprets and processes a user's verbal instructions by analyzing voice data and converting it into text data.

[1578] "Local storage" is a data storage area built into the user's device.

[1579] The present invention relates to a content distribution system for providing "happy false memories" that are not captured by the user, and is implemented as follows.

[1580] Device and Software Configuration

[1581] The hardware used is smartphones, smart glasses, and head-mounted displays (HMDs), such as Google Glass, North Focals, Oculus Rift, and HTC Vive. The software used is Google Cloud Speech-to-Text for voice recognition, generative AI models (e.g., OpenAI's DALL-E, Stable Diffusion) for image generation, and Firebase for the database.

[1582] User Data Collection

[1583] When users first launch the application, they enter their interests and preferences, which can include specific details like "I like to travel" or "I take lots of photos of my pets." This data is sent from their smartphone or device to the cloud and stored in a Firebase database.

[1584] Generate initial album

[1585] The server retrieves the user's interest data from the Firebase database and passes prompts to the generative AI model to generate an image. For example, if a user enters "I like natural scenery," the generative AI model receives the prompt "Generate a photo of a peaceful mountain hike with clear skies and lush greenery." The generated image is constructed as a data packet and sent to the user's device. The received data is stored in local storage on the device, and the image is displayed.

[1586] User Actions

[1587] Users can request the generation of new images using touch or voice input. For example, by saying, "Show me a new beach photo," the voice data is converted to text through Google Cloud Speech-to-Text and sent to the server. The server then provides the generative AI model with a new prompt, which generates a new photo. This process instantly sends the newly generated image to the user's device and displays it.

[1588] Continuous learning and optimization

[1589] The server collects user operation history, records which photos are viewed most frequently, and records which are frequently requested. Based on the collected operation history, the algorithm of the generative AI model is updated and optimized, so that future image generation will be more in line with the user's preferences.

[1590] Specific examples

[1591] For example, if a user makes a voice request such as "Show me new mountain climbing photos," the following prompt sentence is passed to the generative AI model:

[1592] "Generate a photo of a peaceful mountain hike with clear skies and lush greenery"

[1593] Through the above process, users can create and enjoy a variety of photos and videos that suit their tastes in real time.

[1594] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1595] Step 1:

[1596] When a user launches the application for the first time, they enter their hobbies and preferences, such as "I like traveling" or "I take lots of photos of my pets." The entered data is sent from the smartphone or device to the server. The data is then temporarily processed on the user's device and converted into the appropriate format.

[1597] Step 2:

[1598] The server receives the interest and preference data sent by the user and stores it in the Firebase database. The input here is the interest and preference data, and the output is the stored data. The server analyzes the received data and prepares the information necessary for later image generation.

[1599] Step 3:

[1600] The server retrieves interest and preference data from the Firebase database and generates an image by passing a prompt to the generative AI model. For example, a specific prompt such as "Generate a photo of a peaceful mountain hike with clear skies and lush greenery" is provided to the generative AI model. The input is the interest and preference data and the prompt, and the output is the generated image.

[1601] Step 4:

[1602] The generated image is sent as a data packet from the server to the user's device. The input here is the generated image data, and the output is the image data received by the user's terminal. The server converts the image data into an appropriate format and transmits it via a communication protocol.

[1603] Step 5:

[1604] The device stores the image data received from the server in local storage and displays the image, where the input is the received image data and the output is the image visually displayed to the user. The device displays the stored image data in a user interface in an appropriate format.

[1605] Step 6:

[1606] The user requests the generation of a new image using touch or voice input. For example, they might say, "Show me a new beach photo." This voice data is picked up by the device's microphone and converted to text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data.

[1607] Step 7:

[1608] The device sends a voice request, converted into text, to the server, which analyzes the request and generates a new prompt. For example, a prompt such as "Generate a photo of a sunny beach with clear blue water" is passed to a generative AI model. The input is text data, and the output is the new prompt.

[1609] Step 8:

[1610] The server provides the new prompt to the generative AI model and generates a new image. The generated image is then sent back to the user's device. The input is the new prompt and the generative AI model, and the output is the newly generated image.

[1611] Step 9:

[1612] The device saves the newly received image to local storage and immediately displays it to the user. Here, the input is the newly received image data and the output is the image displayed to the user, allowing the user to immediately view the new content they requested.

[1613] Step 10:

[1614] The server continuously collects user operation history, recording which photos are viewed most frequently and which requests are most frequently made. This data is saved as operation history. The input is the user operation data, and the output is the saved operation history.

[1615] Step 11:

[1616] The server updates and optimizes the generative AI model algorithm based on the collected operation history. For example, if there are many requests for a particular situation, the server strengthens the generative algorithm to accommodate that situation. The input is operation history data, and the output is an optimized generative AI model. This process improves image generation from the next time onwards to better match the user's preferences.

[1617] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1618] The present invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that are not actually taken, and is implemented as follows.

[1619] User Data Collection

[1620] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1621] Terminal: Temporarily stores the entered data and sends it to the server, where it is converted into the appropriate format.

[1622] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1623] Introducing the Emotion Engine

[1624] Device: The device acquires the user's facial expressions and tone of voice, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server along with the user data.

[1625] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[1626] Generate initial album

[1627] Server: Generates images using AI technology based on the user's hobby, preference, and emotional data. For example, if a user is recognized as "liking natural scenery" and "relaxed," a photo of a serene natural landscape will be generated.

[1628] Server: Assembles the generated image as a data packet and sends it to the user's device.

[1629] Device: The received image data is stored in local storage and displayed in the Album app for the user to view.

[1630] User interaction and feedback

[1631] Users can browse albums with touch or use voice input to request new images (e.g., "Show me pictures of the beach").

[1632] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[1633] Server: Analyzes the voice request and extracts parameters for generating new images. For example, based on the instruction "Photo of mountain climbing," it generates an appropriate photo and incorporates emotional data.

[1634] Server: Sends the generated new image to the user device.

[1635] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1636] Continuous learning and optimization

[1637] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[1638] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[1639] Specific examples

[1640] First-time setup

[1641] User: Launches the album app for the first time and enters hobby and preference data such as "I like traveling" and "I take a lot of photos of my pets."

[1642] Terminal: Sends input data and the user's initial emotion data to the server.

[1643] Server: Analyzes hobby and preference data and emotional data to generate appropriate travel and pet photos.

[1644] Terminal: Displays the initially generated photo.

[1645] Daily use

[1646] User: While browsing an album, say "Show me new beach photos."

[1647] Device: Sends voice requests and emotion data to the server.

[1648] Server: Generates and sends beach photos based on the user's request and emotion data.

[1649] Device: Generated beach photos are saved to local storage and displayed immediately.

[1650] This system allows users to easily and intuitively experience "happy false memories" that match a specific emotional state.

[1651] The processing flow will be explained below.

[1652] Step 1:

[1653] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1654] Step 2:

[1655] Terminal: Temporarily stores the data entered by the user and sends it to the server, where it is converted into the appropriate format.

[1656] Step 3:

[1657] Server: Stores the received user data in a database and analyzes it. Prepares parameters for generating the initial album based on the user's tastes and preferences.

[1658] Step 4:

[1659] Device: Captures the user's facial expressions and voice tone, and uses the emotion engine to recognize the user's emotions in real time. The recognized emotion data is sent to the server.

[1660] Step 5:

[1661] Server: Analyzes the received emotional data and stores the user's current emotional state in a database.

[1662] Step 6:

[1663] Server: Generates images for the initial album using AI technology based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," it generates photos of tranquil natural scenery.

[1664] Step 7:

[1665] Server: Assembles the generated image as a data packet and handles the communication process to send it to the user's device.

[1666] Step 8:

[1667] Device: The image data received from the server is saved in local storage. The saved images are retrieved and displayed on the initial screen of the album app.

[1668] Step 9:

[1669] User: Browse the album with touch input or use voice input to request the generation of new images, for example, "Show me new beach photos."

[1670] Step 10:

[1671] Device: Detects user touch input and displays the details screen of the selected photo. For voice requests, it uses a speech recognition engine to convert speech into text and send it to the server.

[1672] Step 11:

[1673] Server: Receives voice requests and emotion data, analyzes them, and extracts new image generation parameters. For example, in response to a request for a "beach photo," it generates an appropriate beach photo, taking into account the user's emotion data.

[1674] Step 12:

[1675] Server: Assembles the generated new image as a data packet and sends it to the user's device.

[1676] Step 13:

[1677] On the device: Newly received images from the server are stored in local storage and immediately displayed to the user.

[1678] Step 14:

[1679] Server: Continuously collects user operation history and emotional data, recording which photos are viewed most frequently, what requests are most popular, and changes in user emotions.

[1680] Step 15:

[1681] Server: Updates and optimizes the generation algorithm based on the collected operation history and emotional data, so that future image generation will better match the user's preferences and emotional state.

[1682] By following the above steps, users can easily and intuitively experience a "happy false memory" that matches their specific emotional state.

[1683] Example 2

[1684] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1685] With the recent advancement of digital technology, users are increasingly seeking tools that allow them to freely enjoy their imaginations and memories. However, existing album applications only collect and manage memories that users have clearly experienced, lacking the ability to generate new memories and experiences. Furthermore, they lack the ability to generate and display content that takes into account the user's emotional state, making it difficult to provide an experience that meets the user's individual needs and emotions. Furthermore, with few systems offering intuitive operation through voice input or real-time feedback, users are often forced to use complex operations.

[1686] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting user interest and preference data, means for transmitting user data to the server, means for acquiring user emotion data, means for generating an image based on the received user data and emotion data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting user operation history, means for updating the generation algorithm based on the operation history and emotion data, means for interpreting the user's request using voice recognition and generating a new image, and means for saving the generated image in local storage and immediately displaying it to the user. This makes it possible to easily generate new memories and experiences according to the user's interest and preference and emotional state, and intuitively provide content suitable for each individual user.

[1687] "User" means an individual or legal entity that uses the System.

[1688] "Hobby and preference data" refers to information about the genres and preferences that a user is interested in.

[1689] "Emotional data" refers to information about a user's emotional state obtained from facial expressions, tone of voice, etc.

[1690] "Server" refers to a computer system that processes, stores, and analyzes data submitted by users.

[1691] "Terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user.

[1692] "Received User Data" refers to information about a User that the Server receives from a Terminal.

[1693] "Image generation" refers to the process of creating new images using AI technology.

[1694] "Display means" refers to the method or technology that allows the generated image to be visually viewed on the user's device.

[1695] "Operation history" refers to the record of various operations performed by a user using the system.

[1696] "Generation algorithm" refers to the calculation procedure or model for generating an image based on received data.

[1697] A "voice recognition engine" refers to software that analyzes a user's voice and converts it into text.

[1698] "Local storage" refers to the data storage area within the user's device.

[1699] This invention is a system that combines an emotion engine with an album app that provides users with "happy false memories" that they have not taken. This system provides users with newly generated images using AI technology based on the user's taste data and emotion data.

[1700] Hardware and software used

[1701] Device: A device that is directly operated by the user, such as a smartphone, tablet, or personal computer, is used. The device is equipped with a camera and microphone, which are used to capture the user's facial expressions and voice.

[1702] Server: Use a high-performance cloud server or data center. The server performs central processing such as data storage, analysis, and image generation. Specific examples include servers provided by major cloud service providers (e.g., Amazon Web Services, Google Cloud Platform).

[1703] Software: The main software components of this system include an emotion engine, a speech recognition engine, and a generative AI model. We use OpenCV for the emotion engine and Google Cloud Speech-to-Text for the speech recognition engine. We use generative AI models (e.g., DALL·E, Stable Diffusion) for image generation.

[1704] Data flow and processing

[1705] 1. User data input

[1706] When a user launches the album app for the first time, they enter their name and hobbies (e.g., "travel," "nature scenery," "pet photos," etc.). The device temporarily stores this data, converts it into an appropriate format, and sends it to the server.

[1707] 2. Analysis and saving of initial setting data

[1708] The server analyzes the received user data and stores it in a database. During this analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[1709] 3. Acquiring and sending emotion data

[1710] The device's camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This emotion data is sent to a server, where it is analyzed and saved.

[1711] 4. Initial album generation

[1712] The server generates images using a generative AI model based on the user's hobby, preference, and emotional data. For example, if the user is recognized as "liking natural scenery" and "relaxed," a photo of a calm natural scenery will be generated. An example of a prompt sentence is "relaxing natural scenery."

[1713] 5. Sending and displaying images

[1714] The server then assembles the generated images into a data packet and sends it to the user's device, which receives it, stores it in local storage, and displays it in the album app.

[1715] User interaction and feedback

[1716] Users can browse the album with touch gestures and use voice input to request the generation of new images (e.g., "Show me new beach photos.") The device receives the voice request, converts it into text using a speech recognition engine, and sends it to the server.

[1717] The server analyzes the voice request and extracts parameters for generating a new image. For example, a prompt such as "beach photo" is input into the generative AI model, which then generates an image incorporating the user's emotional data. The image is then sent to the device, stored in local storage, and instantly displayed to the user.

[1718] Continuous learning and optimization

[1719] The server continuously collects the user's operation history and emotional data, and updates and optimizes the generation algorithm based on this data, so that the next image generation will be more in line with the user's tastes, preferences, and emotional state.

[1720] As described above, the present invention is a system that provides "happy false memories" that take into account the individual preferences and emotions of the user, and can provide the user with a new experience in an intuitive and simple manner.

[1721] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1722] Step 1: User Data Input

[1723] User: When launching the Album app for the first time, enter your name and hobbies and preferences (e.g., "Travel," "Nature Scenery," "Pet Photos") on the initial setup screen.

[1724] Input: User name, hobbies and preferences

[1725] Output: Formatted user data (e.g., JSON format)

[1726] Specific behavior: The user enters their name and hobbies and preferences into the text boxes and check boxes, then presses the submit button.

[1727] Step 2: Send data

[1728] Terminal: Temporarily stores the input data, converts it into an appropriate format (e.g., JSON format), and sends it to the server.

[1729] Input: User data (name, hobbies, preferences)

[1730] Output: User data sent to the server

[1731] Specific behavior: After format conversion, the data is sent to the server using an HTTP request.

[1732] Step 3: Analyze and save the initial setup data

[1733] Server: Analyzes the received user data and stores it in a database. During the analysis, parameters for generating the initial album are extracted based on the user's tastes and preferences.

[1734] Input: Received user data

[1735] Output: Parsed user data stored in a database

[1736] Specific operation: Apply data analysis algorithms to generate parameters based on the user's preferences. Store the analysis results in a database.

[1737] Step 4: Obtaining and sending emotion data

[1738] Device: The camera and microphone capture the user's facial expressions and voice tone, and the emotion engine recognizes emotions in real time. This data is sent to the server.

[1739] Input: User's facial expression data, voice tone data

[1740] Output: Emotion data sent to the server

[1741] Specific operation: The camera captures faces in real time and records voice input. These data are analyzed by the emotion engine, and emotion data is extracted and sent to the server.

[1742] Step 5: Generate the initial album

[1743] Server: Generates images using a generative AI model based on the user's taste and emotion data.

[1744] Input: Interest data, emotion data, prompt sentence (e.g., "Relaxing nature scenery")

[1745] Output: Generated image data

[1746] Specific operation: A prompt sentence and emotion data are input to the generative AI model, and the AI ​​generates an image. The generated image is converted into a data packet.

[1747] Step 6: Send and view images

[1748] Server: Assembles the generated image as a data packet and sends it to the user's device.

[1749] Device: The received image data is stored in local storage and displayed to the user in the album app.

[1750] Input: Generated image data

[1751] Output: Image displayed in the album app

[1752] Specific behavior: Receive image data from the server, save it to local storage, and display the image in the app's user interface.

[1753] Step 7: Browse albums and request new images

[1754] User: Browse albums with touch and use voice input to request new images (e.g., "Show me new beach photos").

[1755] Input: Touch, voice requests

[1756] Output: New image request data

[1757] Specific actions: Swipe to change images or use the microphone to input voice commands.

[1758] Step 8: Sending and analyzing audio data

[1759] Terminal: Receives voice requests, converts them into text using a speech recognition engine, and sends them to the server.

[1760] Input: Voice request

[1761] Output: The request text sent to the server

[1762] Specific operation: Converts speech to text and sends the text data to the server.

[1763] Step 9: Generate and send a new image

[1764] Server: Analyzes the voice request, extracts parameters for generating a new image, generates the image using a generative AI model, and sends it to the user's device.

[1765] Input: Voice request text, emotion data

[1766] Output: The new image data generated.

[1767] Specific operation: Analyzes the voice request and generates a prompt (e.g., "New beach photo"). Generates an image using a generative AI model, converts it into a data packet, and sends it to the device.

[1768] Step 10: Displaying the new image

[1769] On the device: The new image is generated and saved to local storage, and is immediately displayed to the user.

[1770] Input: Newly generated image data

[1771] Output: The new image displayed in the Album app

[1772] Specific behavior: Receives image data from the server, saves it to local storage, and displays the image in the album app's user interface.

[1773] Step 11: Continue learning and optimization

[1774] Server: Continuously collects user operation history and emotion data, and updates and optimizes the generation algorithm.

[1775] Input: Operation history, emotion data

[1776] Output: Updated and optimized generation algorithm

[1777] Specific operation: Obtain operation history and emotion data from the database. Based on this data, the generation algorithm is retrained and optimized.

[1778] (Application example 2)

[1779] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1780] While existing virtual stores recommend products based on users' preferences, they do not use real-time emotional data to recommend products that reflect the user's current emotional state. As a result, the accuracy of product recommendations to users is low, limiting the improvement of the user experience. In addition, systems linked to voice recognition have not been fully utilized when users make specific requests, making them difficult to use intuitively.

[1781] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1782] In this invention, the server includes means for inputting user interest and preference data, means for transmitting the user data to the server, means for generating an image based on the received data, means for transmitting the generated image to the user's terminal, means for displaying the image transmitted to the terminal, means for collecting the user's operation history, means for updating the generation algorithm based on the operation history, means for acquiring and analyzing the user's emotional data in real time, means for recommending products based on the acquired emotional data, and means for displaying the products on the terminal. This enables personalized product recommendations based on the user's emotional state in real time, improving the user experience and enabling intuitive operation.

[1783] "Hobbies and Preference Data" is information that indicates a user's interests, favorite activities, and hobbies.

[1784] "Emotional data" is data that represents the emotional state of a user, obtained from the user's facial expression, tone of voice, etc.

[1785] An "image generation means" is an algorithm or software that creates an image based on the data received.

[1786] The "product recommendation means" is a system that suggests products suitable for the user based on emotional data and hobby and preference data.

[1787] A "server" is a remote computer or system that has functions such as receiving, analyzing, and storing data, generating images, and recommending products.

[1788] "Terminal" means a device that can be directly operated by a user, including a smartphone, smart glasses, tablet, etc.

[1789] "Operation history" refers to historical data of operations and actions taken by a user when using the system.

[1790] "Speech recognition" is a technology that converts a user's voice into text data and understands the instructions based on that.

[1791] "Real-time" refers to data acquisition, analysis, and response occurring immediately and without delay.

[1792] This invention is a virtual store system that allows users to acquire their own emotional data in real time using their smart devices and recommends personalized products based on that data. To implement this invention, the following means are required.

[1793] First, the user puts on the smart glasses and starts the application. At startup, the user enters their name and hobby / preference data (e.g., "outdoor goods" or "casual fashion"). This data is temporarily stored in the smart glasses and sent to the server.

[1794] Next, the server prepares parameters for generating the initial album based on the received user's tastes and preferences. It also acquires the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine. The analyzed emotion data is sent to the server and stored in a database.

[1795] The server combines the user's taste and preference data with real-time emotional data to generate an image using a generative AI model. For example, if the user is in a "relaxed" state and prefers "outdoor goods," an image of outdoor goods with a relaxed atmosphere will be generated. The generated image and product information are organized into a data packet and sent to the user's smart glasses.

[1796] The smart glasses display the received image data and product information. The user can use voice requests to instruct the server to generate additional images. The voice requests are converted into text by a speech recognition engine, and the data is sent to the server. The server analyzes the voice requests and makes appropriate product recommendations.

[1797] The software components of this application include:

[1798] OpenCV: Used to analyze user facial expressions.

[1799] EmotionEngine: A custom AI model that identifies user emotions in real time.

[1800] ProductRecommender: An algorithm that provides personalized product recommendations based on user sentiment data.

[1801] ARDisplay: An augmented reality library that displays products virtually in the field of view of smart glasses.

[1802] As a concrete example, when a user enters a shopping mall and shows a relaxed expression, the emotion engine recognizes the state as "relaxed." Based on this, the server recommends "casual fashion items" and displays them on the smart glasses' display. When the user makes a voice request such as "Show me a new casual shirt," the voice is converted into text data and sent to the server. Based on this request, the server generates an image of an appropriate shirt and sends it back to the smart glasses.

[1803] Example prompt sentence:

[1804] Prompt: Recommend products suitable for the user when they are relaxed. The user is a man in his 30s who likes outdoor gear.

[1805] Example output: Recommendations include "Casual T-shirt", "Outdoor backpack", and "Sports sunglasses".

[1806] This enables personalized product suggestions based on user sentiment in real time, improving the user experience and providing an intuitive and convenient shopping experience.

[1807] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1808] Step 1:

[1809] Input of user's hobby and preference data

[1810] How it works: The user puts on the smart glasses and launches the application. On first launch, the user enters their name and preferences (e.g., "outdoor gear" or "casual fashion").

[1811] Input: User's basic information and hobbies and preferences

[1812] Data processing: The input data is temporarily stored in the smart glasses and converted into the required format.

[1813] Output: Formatted user data

[1814] Step 2:

[1815] Sending user data

[1816] How it works: The smart glasses send formatted user data to the server.

[1817] Input: Formatted user data

[1818] Data processing: None

[1819] Output: User data sent to the server

[1820] Step 3:

[1821] First, prepare to generate an album based on hobby and preference data.

[1822] Operation: The server analyzes the received user data and prepares parameters for generating the initial album based on the user's tastes and preferences.

[1823] Input: User data sent to the server

[1824] Data Calculation: Analyzing user data

[1825] Output: Parameters for initial album generation

[1826] Step 4:

[1827] Emotion data collection and analysis

[1828] How it works: The smart glasses capture the user's facial expressions and voice tone in real time, analyze these data with the emotion engine, and send the emotion data to the server.

[1829] Input: User facial expressions and voice tone

[1830] Data Computation: Emotion Analysis with Emotion Engine

[1831] Output: Emotion data

[1832] Step 5:

[1833] Emotion data storage and analysis

[1834] How it works: The server analyzes the received emotion data and stores it in a database.

[1835] Input: Emotion data

[1836] Data Computing: Sentiment Data Analysis and Storage

[1837] Output: Emotion data stored in a database

[1838] Step 6:

[1839] Image and product recommendation generation

[1840] How it works: By combining user preference data with real-time emotional data, a generative AI model is used to generate images and prepare product information.

[1841] Input: Parameters for initial album generation, emotion data

[1842] Data processing: Image generation using generative AI models, creation of product information

[1843] Output: Generated image data and product information

[1844] Step 7:

[1845] Submitting images and product information

[1846] Operation: The server assembles the generated image data and product information into a data packet and sends it to the user's smart glasses.

[1847] Input: Generated image data and product information

[1848] Data processing: Data packet construction

[1849] Output: Image data and product information sent to smart glasses

[1850] Step 8:

[1851] Display of images and product information

[1852] Operation: The smart glasses display the received image data and product information.

[1853] Input: Image data and product information sent to smart glasses

[1854] Data processing: None

[1855] Output: Image and product information displayed to the user

[1856] Step 9:

[1857] Processing user requests

[1858] How it works: A user uses a voice request to request additional images or product information. The smart glasses convert the voice request into text and send it to the server.

[1859] Input: User's voice request

[1860] Data calculation: Text conversion using a speech recognition engine

[1861] Output: Text request sent to the server

[1862] Step 10:

[1863] Generate and submit new images and products

[1864] How it works: The server analyzes the voice request, generates the appropriate images and products, and sends them to the smart glasses.

[1865] Input: The text request sent to the server

[1866] Data processing: Generative AI models generate new images and suggest products, and create data packets

[1867] Output: New image and product information sent to the smart glasses

[1868] Step 11:

[1869] New images and product displays

[1870] How it works: The smart glasses display the new images and product information they receive.

[1871] Input: New image and product information sent to the smart glasses

[1872] Data processing: None

[1873] Output: The new image and product information displayed to the user

[1874] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1875] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1876] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1877] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1878] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1879] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1880] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1881] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1882] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1883] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1884] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1885] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1886] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1887] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1888] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1889] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1890] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1891] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1892] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1893] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1894] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1895] The following is further disclosed regarding the above embodiment.

[1896] (Claim 1)

[1897] A means for inputting user's interest and preference data;

[1898] means for transmitting user data to a server;

[1899] means for generating an image based on the received data;

[1900] means for transmitting the generated image to a user's terminal;

[1901] means for displaying the image transmitted to the terminal;

[1902] A means for collecting user operation history;

[1903] A means for updating the generation algorithm based on the operation history;

[1904] A system including:

[1905] (Claim 2)

[1906] 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition to generate an image.

[1907] (Claim 3)

[1908] 10. The system of claim 1, further comprising means for storing the generated image in local storage and immediately displaying it to the user.

[1909] "Example 1"

[1910] (Claim 1)

[1911] a means for inputting user preference information;

[1912] means for transmitting preference information to a server;

[1913] means for generating an image based on the received preference information;

[1914] means for transmitting the generated image to a user device;

[1915] means for displaying the image transmitted to the device;

[1916] A means for collecting user operation history;

[1917] A means for updating the generation algorithm based on the operation history;

[1918] A system including:

[1919] (Claim 2)

[1920] 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition to generate an image.

[1921] (Claim 3)

[1922] 10. The system of claim 1, further comprising means for storing the generated image in local storage and for immediate display to the user.

[1923] "Application Example 1"

[1924] (Claim 1)

[1925] A means for inputting user's interest and preference data;

[1926] means for transmitting user data to a server;

[1927] a means for generating an image using a generative AI model based on the received data;

[1928] means for transmitting the generated image to a user's terminal;

[1929] means for displaying the image transmitted to the terminal;

[1930] A means for collecting user operation history;

[1931] A means for updating the generation algorithm based on the operation history using prompt statements;

[1932] A system including:

[1933] (Claim 2)

[1934] 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition technology to generate an image.

[1935] (Claim 3)

[1936] 10. The system of claim 1, further comprising means for storing the generated image in local storage and immediately displaying it to the user.

[1937] "Example 2: Combining Emotion Engines"

[1938] (Claim 1)

[1939] A means for inputting user's interest and preference data;

[1940] means for transmitting user data to a server;

[1941] A means for acquiring user emotion data;

[1942] means for generating an image based on the received user data and emotion data;

[1943] means for transmitting the generated image to a user's terminal;

[1944] means for displaying the image transmitted to the terminal;

[1945] A means for collecting user operation history;

[1946] A means for updating the generation algorithm based on the operation history and emotion data;

[1947] A system including:

[1948] (Claim 2)

[1949] 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition to generate an image.

[1950] (Claim 3)

[1951] 10. The system of claim 1, further comprising means for storing the generated image in local storage and immediately displaying it to the user.

[1952] "Application example 2 when combining emotion engines"

[1953] (Claim 1)

[1954] A means for inputting user's interest and preference data;

[1955] means for transmitting user data to a server;

[1956] means for generating an image based on the received data;

[1957] means for transmitting the generated image to a user's terminal;

[1958] means for displaying the image transmitted to the terminal;

[1959] A means for collecting user operation history;

[1960] A means for updating the generation algorithm based on the operation history;

[1961] A means of acquiring and analyzing user emotional data in real time,

[1962] A means for recommending products based on the acquired emotion data;

[1963] a means for displaying the product on the terminal;

[1964] A system including:

[1965] (Claim 2)

[1966] 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition to generate an image.

[1967] (Claim 3)

[1968] 10. The system of claim 1, further comprising means for storing the generated images and recommended products in a local storage and for immediately displaying them to the user. [Explanation of symbols]

[1969] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for inputting user interest and preference data; means for transmitting user data to a server; means for generating an image based on the received data; means for transmitting the generated image to a user's terminal; means for displaying the image transmitted to the terminal; A means for collecting user operation history; A means for updating the generation algorithm based on the operation history; A system including:

2. 10. The system of claim 1, further comprising means for interpreting a user request using voice recognition to generate an image.

3. The system of claim 1 further comprising means for storing the generated image in local storage and for immediately displaying it to the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A