System

The system enhances album browsing by generating and playing music and scent data based on selected photos and entered episodes, allowing users to relive memories in a richer, sensory way.

JP2026015017APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116491
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing album browsing experiences lack the ability to recreate memories in a richer, sensory way, as they primarily rely on visual information and do not incorporate elements like music and scent, making it difficult to relive past events vividly.

Method used

A system that allows users to select photos and input episodes, generating related music data, scent data, and detailed episode text, which are played and diffused to recreate memories in a three-dimensional manner, incorporating location and time information from photo metadata.

Benefits of technology

Enables users to relive memories through a combination of visual, auditory, and olfactory sensory elements, providing a more immersive and engaging experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015017000001_ABST
    Figure 2026015017000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for a user to select a photo related to a past memory; means for the user to input an episode related to the selected photo; means for generating related music data, scent data, and episode text based on the selected photo and the input episode; means for transmitting the generated music data, scent data, and episode text to a terminal of the user; and means for playing the transmitted music data, spraying the scent data by a diffuser, and displaying the episode text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In the past, album browsing experiences simply allowed users to look at photos, making it difficult to vividly relive memories. Furthermore, handwritten diaries and digital notebooks lack the ability to complement sensory elements (such as music and scent) other than visual information. This prevented users from re-experiencing past events in a richer way, leaving memories feeling monotonous overall. [Means for solving the problem]

[0005] The present invention provides a means for users to select photos related to past memories and input episodes related to those photos. It also provides a means for generating related music data, scent data, and detailed episode text based on the selected photos and input episodes. This generated content is sent to the user's device, and the device plays music, sprays scents using a diffuser, and displays the episode text, allowing the user to re-experience memories in a more sensory-rich and three-dimensional way. The system also includes a means for extracting location and time information from photo metadata and generating related detailed information through analysis of the user's episodes, thereby enabling more accurate memory reproduction.

[0006] A "user" is someone who uses the system to re-experience past memories.

[0007] A "photo" is an image selected by the user that visually records past events or memories.

[0008] An "episode" is textual information that users enter in relation to a photo, describing details of an event or a memory.

[0009] "Music Data" refers to music files and audio content generated to complement the user's memories.

[0010] "Scent data" is digital scent information used to recreate the user's memories.

[0011] "Episode text" is text that includes detailed explanations and supplementary information generated based on the episode entered by the user.

[0012] "Means" refers to a device or method used to achieve a specific function or purpose.

[0013] "Terminal" means the electronic device used by a User to operate the System and display and play Content.

[0014] A "server" is a computer system that receives user requests, analyzes data, generates content, and sends the results to the user's terminal.

[0015] "Metadata" is additional information such as location and time information that accompanies a photo.

[0016] A "diffuser" is a device that diffuses a specific scent based on scent data. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention relates to a system that generates music data, scent data, and episode text by having a user select a photo related to a past memory and input an episode related to that photo, and provides these to a terminal operated by the user. The specific operation and program processing of the system are described below.

[0039] Server Processing

[0040] 1. Request acceptance

[0041] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in the summer of 2022" and sends a request to recreate those memories, the server receives the request.

[0042] 2. Data Analysis

[0043] The server analyzes the received request data. At this time, it extracts the metadata of the photo (location information, time information, etc.) and also analyzes the episode entered by the user. For example, it analyzes the episode "I had a beach party with my friends on this day and it was so much fun" and extracts related keywords.

[0044] 3. Content Generation

[0045] Based on the analyzed data, the server generates music data, scent data, and detailed episode text that match the selected photo and episode. For example, based on the keywords "beach," "summer," and "party," the server generates a surf rock playlist that matches a summer beach party, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[0046] 4. Content Submission

[0047] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[0048] Terminal handling

[0049] 1. Interface provision

[0050] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select a photo they want to relive and enter an episode.

[0051] 2. Submit a request

[0052] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[0053] 3. Receiving and Displaying Content

[0054] The device receives the content package sent from the server, checks the received content, plays the received music data, sprays a scent based on the scent data using the diffuser, and displays the episode text on the screen.

[0055] User Actions

[0056] 1. Select a photo and enter an episode

[0057] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[0058] 2. Re-experiencing memories

[0059] Users can enjoy the photos and detailed story text while listening to music and smelling the scent emanating from the diffuser through their device. For example, while reading the detailed story text with surf rock music playing in the background and the scent of the beach diffusing, they can vividly relive memories of past beach parties.

[0060] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive their past memories in a richer way. By combining music, scent, and detailed text, the system recreates the entire memory in three dimensions, providing users with a new experience.

[0061] The processing flow will be explained below.

[0062] Step 1: Select a photo

[0063] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[0064] Step 2: Episode Input

[0065] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[0066] Step 3: Create a request

[0067] The terminal packages the selected photo and the input episode, and creates request data to be sent to the server.

[0068] Step 4: Submitting the request

[0069] The terminal transmits the created request data to the server.

[0070] Step 5: Receiving the request

[0071] The server receives the request data sent from the terminal.

[0072] Step 6: Data analysis

[0073] The server analyzes the received request data, extracts location and time information from the photo metadata, and analyzes the episodes entered by the user for keywords.

[0074] Step 7: Content Generation

[0075] The server generates music data, scent data, and episode text based on the analyzed data. For example, based on keywords such as "beach," "summer," and "party," the server generates a surf rock playlist, beach scent data, and detailed episode text.

[0076] Step 8: Submit content

[0077] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[0078] Step 9: Receive the package

[0079] The terminal receives the content package sent from the server and checks the received content.

[0080] Step 10: Prepare your content

[0081] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0082] Step 11: Play content

[0083] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[0084] Step 12: Re-experience the memory

[0085] Users re-experience past memories in three dimensions through displayed photos, played music, scents, and episode text.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] Conventional photo album viewing systems only allow users to visually re-experience photos, and lack the means to recreate memories using senses other than sight. In particular, there is a demand for systems that allow users to re-experience past memories in a more three-dimensional way by combining sensory elements such as music and scent. A method to meet this demand is needed.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for generating associated sound data, scent data, and episode text based on the selected image and the input episode, means for transmitting the generated sound data, scent data, and episode text to the user's output device, and means for playing the transmitted sound data, spraying the scent data with a spraying device, and displaying the episode text, thereby enabling the user to re-experience memories not only through images but also through music and scents.

[0091] A "user" is someone who uses the system to re-experience past memories.

[0092] "Images" are media containing visual information that a user selects to associate with past memories.

[0093] An "episode" is descriptive information about memories or experiences that a user enters in relation to a selected image.

[0094] "Sound data" is data that includes information about music or other sounds associated with a user's memories.

[0095] "Scent data" is data that includes information about scents and fragrances associated with a user's memories.

[0096] "Episode text" refers to an episode entered by a user and complementary text information generated based on the episode.

[0097] A "server" is a central processing unit that analyzes information sent by a user, generates the necessary content, and sends it to the user's terminal.

[0098] An "output device" is a device that allows the user to re-experience the generated audio data, scent data, and episode text.

[0099] A "spraying device" is a device for diffusing a scent based on scent data.

[0100] The present invention is a system that generates sound data, scent data, and episode text by allowing a user to select an image related to a past memory and input an episode related to that image, and provides these to an output device operated by the user. This system operates in cooperation with a server, a terminal, and the user.

[0101] Server Operation

[0102] Request reception

[0103] The server uses Apache HTTP server software to accept requests sent from users' devices. Specifically, an endpoint is set up using the Python Flask framework, and the user selects "images taken at the beach in the summer of 2022" and submits the episode to receive the request.

[0104] Data analysis

[0105] The server parses the data received by the Flask application, using the Pandas library to extract image metadata (location, time, etc.), and the Natural Language Toolkit (NLTK) to parse the user's episodes, extracting keywords such as "beach," "summer," and "party."

[0106] Content Generation

[0107] The server generates the necessary content based on the analysis results. It uses the Spotify API to generate audio data, a pre-built scent database to generate scent data, and GPT-3 to generate episode text. For example, it generates a surf rock playlist, beach scents, and detailed episode text that match the themes of "beach," "summer," and "party."

[0108] Content Submission

[0109] The server sends the generated audio data, scent data, and episode text to the user's output device, where the data is compressed in gzip format and returned as an HTTP response.

[0110] Device behavior

[0111] Interface provided

[0112] The device provides a React-based web interface that displays file selection buttons and text input fields so users can select images from the album screen and enter episodes.

[0113] Send request

[0114] The device packages the data entered by the user into JSON format and sends it to the server. The selected image metadata and episode text are combined into a single JSON object and sent as an HTTP POST request to the Flask endpoint.

[0115] Receiving and displaying content

[0116] The device receives the content returned from the server. The received data is displayed on the screen using React, the sound data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed in a text area.

[0117] User Actions

[0118] Selecting photos and entering episodes

[0119] Users can select images related to past memories from the device's album display screen. Specifically, they can select "Images taken at the beach in the summer of 2022" and enter an episode such as "I had a beach party with friends on this day and it was so much fun."

[0120] Re-experiencing memories

[0121] Users re-experience the content through their device. While the audio data is played and the scent is released from the diffuser, they read the detailed episode text, vividly recreating past memories. For example, they can re-read the detailed episode text while surf rock music plays in the background and the scent of the beach fills the air.

[0122] Examples and prompts

[0123] Specific examples

[0124] Example of input photo: "Photo taken at the beach in summer 2022"

[0125] Example episode: "I had a beach party with my friends that day and it was so much fun. We surfed and had a great time."

[0126] Prompt Sentence Examples

[0127] Photo episode generation:

[0128] Generate photo episodes.

[0129] Photo content: Partying with friends on the beach

[0130] Episode summary: I had a beach party with my friends that day and it was so much fun. We surfed and had a great time.

[0131] Music Playlist Generation:

[0132] Create a music playlist that will go well with your fun beach party.

[0133] Keywords: beach, summer, party

[0134] Scent data generation:

[0135] Provide the perfect scent data for your beach memories.

[0136] Keywords: beach, summer, party

[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0138] Step 1: Request acceptance

[0139] The server uses Apache HTTP server software to accept requests from users. The user selects "Photos taken at the beach in the summer of 2022" and submits the request by sending that memory as an episode. This request is sent to the server as JSON-formatted data containing an image file and an episode description. The selected image file and episode description are provided as input, and the request content is retained as output.

[0140] Step 2: Data analysis

[0141] The server uses the Flask framework to analyze the received request data. It extracts image data metadata (e.g., location and time information) using the Pandas library, and analyzes the episode sentences using NLTK. Specifically, it obtains the photo location and date and time from the image's EXIF ​​information, and extracts keywords such as "beach," "summer," and "party" from the episode sentences. The input is the image data and the episode sentences, and the output is the extracted metadata and keywords.

[0142] Step 3: Content generation

[0143] The server generates related audio data, scent data, and episode text based on the parsed metadata and keywords. It uses the Spotify API to generate a playlist based on keywords, selects appropriate scents from the scent database, and generates episode text using GPT-3. For example, for the keywords "beach," "summer," and "party," it generates surf rock music, beach scents, and detailed episode text that match the keywords. The input is the parsed metadata and keywords, and the output is audio data, scent data, and episode text.

[0144] Step 4: Submit content

[0145] The server sends the generated acoustic data, scent data, and episode text to the user's device. The data is compressed in gzip format for efficient transfer and returned as an HTTP response. The input is the data obtained in the content generation step, and the output is the data sent to the user's device in compressed form.

[0146] Step 5: Provide an interface

[0147] The device provides a web interface built using React. It displays a screen where users can select an image related to a past memory and enter an episode about it. For example, when a user selects a photo on the album screen, a thumbnail preview is displayed and a text area for entering an episode is displayed. The input is the user's operation, and the output is the screen display.

[0148] Step 6: Submitting the request

[0149] The device packages the data entered by the user in JSON format and sends it to the server. Specifically, it combines the metadata of the selected image and the episode text into a single JSON object and sends it as an HTTP POST request. The input is the image data and episode text entered by the user, and the output is the request data sent to the server.

[0150] Step 7: Receiving and displaying content

[0151] The device receives the content sent from the server and displays it on the screen using React. The audio data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed on the screen. For example, when you press the play button, surf rock music plays, the beach scent rises from the diffuser, and the detailed episode text is displayed. The input is the data received from the server, and the output is the display of the user interface and playback operation based on that data.

[0152] Step 8: Select photos and enter episodes

[0153] The user selects an image related to a past memory from the album display screen of the device and enters an episode about that moment. For example, the user selects "an image taken at the beach in the summer of 2022" and enters an episode such as "I had a beach party with friends that day and it was a lot of fun." The input is a past image and an episode description, and the output is the request data sent to the server.

[0154] Step 9: Re-experience the memory

[0155] The user re-experiences the content through the device. The audio data sent from the server is played, and the diffuser emits a scent, while the user reads the episode text and relives past memories. For example, the user re-reads a detailed episode while listening to surf rock music and smelling the scent of the beach. The input is the content displayed on the device, and the output is the sensory memory experience that the user re-experiences.

[0156] (Application example 1)

[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0158] Conventional album viewing systems only allowed users to relive past memories using photos, making it difficult to relive memories that involved sensory elements. Furthermore, photos and text alone did not allow users to relive memories through various senses, such as the atmosphere, scent, and music of the place. This limited the user's experience, making it difficult to relive memories in a richer way.

[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0160] In this invention, the server includes means for a user to select an image related to a past record, means for the user to input an event related to the selected image, means for generating related music data, fragrance data, and event text based on the selected image and the input event, means for transmitting the generated music data, fragrance data, and event text to the user's terminal, means for playing the transmitted music data, diffusing the fragrance data by an atomizer, and displaying the event text, and means for reproducing sensory elements based on the image and event selected by the user, thereby enabling the user to re-experience past memories along with the music, fragrance, and detailed episode text.

[0161] "Past Records" refers to image and photo files saved by users relating to previous events.

[0162] "Image" means digital photographic or pictorial data stored on a User's device that visually represents a past event.

[0163] "Event" refers to textual information that describes a specific experience or situation related to the image the user selected.

[0164] "Music Data" refers to digital music files that provide an acoustic ambience generated based on selected images and input events.

[0165] "Scent data" refers to digital data containing scent information that is generated based on a selected image and an input event.

[0166] "Event text" refers to text data containing detailed event information entered by the user.

[0167] "Terminal" refers to an electronic device operated by a user, such as a smartphone, tablet, or personal computer.

[0168] "Playing music data" refers to the act of outputting digital music files as sound on the user's device.

[0169] "Spreading fragrance data using an atomizer" refers to the act of spreading a fragrance using a device that physically releases fragrance information as digital data as a scent.

[0170] "Recreating sensory elements" refers to the act of allowing the user to experience the atmosphere of a location through various senses, such as music, scent, and text, based on the selected image and input event.

[0171] This invention relates to a system that allows a user to select an image related to a past record and input an event related to that image, thereby generating music data, aroma data, and event text, and providing these to the user's terminal.

[0172] System program description

[0173] The system operates as follows.

[0174] 1. User Interface

[0175] The device provides an interface for users to select historical images and enter events, including a photo selection screen and a text field for entering events.

[0176] 2. Submit a request

[0177] The device sends the image selected by the user and the event entered to the server. The request data includes the image file and the text entered by the user.

[0178] 3. Data analysis and content generation

[0179] The server analyzes the received request data, extracts image metadata (location information, time information, etc.), and generates related keywords based on the event entered by the user. Based on this, music data, scent data, and event text are generated.

[0180] 4. Content Submission

[0181] The server sends the generated music data, aroma data, and event text to the user's terminal, where the data is appropriately compressed and transmitted.

[0182] 5. Playing and Displaying Content

[0183] The terminal plays the received music data, controls the atomizer based on the scent data, and displays the event text on the screen.

[0184] Hardware and software used

[0185] Hardware:

[0186] Smartphones: Providing the user interface and displaying content

[0187] Smart glasses and head-mounted displays: Providing user interfaces and immersive experiences

[0188] Diffuser: Diffuses fragrance based on aroma data

[0189] software:

[0190] Python: Server-side data analysis and content generation

[0191] Requests module: Data communication between the server and the terminal

[0192] Backend (e.g., Flask or Django): Manages request processing and data generation

[0193] Front-end (e.g. React or Vue.js): Building the user interface

[0194] Specific operation example

[0195] A user opens the app on their smartphone, selects an image from a past family trip, and enters an event such as, "On this day, my family visited a theme park and had a wonderful day." The server receives this request, analyzes the image metadata and the event, and generates related fun theme music, popcorn scent data, and detailed episode text for that time. This data is sent to the user's device, which plays music, emits the popcorn scent from a diffuser, and displays the event text on the screen.

[0196] Example prompt sentence:

[0197] Photo: Family Travel Theme Park.jpg

[0198] Episode: On this day, my family and I went to a theme park and had a really fun day.

[0199] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive past memories in a richer way.

[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0201] Step 1:

[0202] The user opens the app on their smartphone, selects an image related to a past record from the album screen, and enters the event related to the image in the text field. User input: Image file and event text. Output: User-selected image and event text.

[0203] Step 2:

[0204] The terminal packages the image and event entered by the user and sends it to the server. Input: Image and event text selected by the user. Data processing: Packages the image file and text. Output: Request data sent to the server.

[0205] Step 3:

[0206] The server analyzes the received request data, extracts image metadata (location, time, etc.), and generates related keywords from the event text entered by the user. Input: Request data. Data operation: Metadata extraction and keyword generation. Output: Identified metadata and keywords.

[0207] Step 4:

[0208] The server generates related music data, aroma data, and detailed event text based on the identified metadata and keywords. Input: Metadata and keywords. Data operation: Generation of music data, aroma data, and event text. Output: Generated music data, aroma data, and event text.

[0209] Step 5:

[0210] The server appropriately compresses the generated music data, aroma data, and event text and sends them to the user's terminal. Input: Generated music data, aroma data, and event text. Data processing: Data compression. Output: Data package sent to the terminal.

[0211] Step 6:

[0212] The terminal unpacks the data package received from the server, plays the music data, controls the sprayer based on the scent data, and displays the event text on the screen. Input: Data package sent from the server. Data processing: Unpacking the data. Output: Music to be played, scent to be diffused, text to be displayed.

[0213] Step 7:

[0214] The user listens to music provided through the device, smells the scent, and reads the detailed event text displayed on the screen, reliving past memories with all five senses. Input: Played music, diffused scent, displayed text. Output: Re-experiencing the user's memories.

[0215] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0216] This invention relates to an album system that allows users to re-experience past memories in a richer, more sensory way by incorporating an emotion engine. This system recognizes the user's emotions and can customize music data, scent data, and episode text based on those emotions. The specific operation and program processing of this system are described below.

[0217] Server Processing

[0218] 1. Request acceptance

[0219] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in summer 2022" and sends a request to recreate those memories, the server receives the request.

[0220] 2. Data Analysis

[0221] The server analyzes the received request data. During this process, it extracts the photo's metadata (location information, time information, etc.) and analyzes the episode entered by the user. It also uses an emotion engine to recognize the user's emotions and adds this information to the analysis. For example, the emotion "joy" is recognized for the episode "I had a beach party with my friends that day and had a lot of fun."

[0222] 3. Content Generation

[0223] Based on the analyzed data, the server generates music data, scent data, and episode text that match the selected photo and episode. The generated content is then customized based on the emotional information recognized by the emotion engine. For example, keywords such as "beach," "summer," and "party" are combined with the emotional information of "joy" to generate a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[0224] 4. Content Submission

[0225] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[0226] Terminal handling

[0227] 1. Interface provision

[0228] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select the photo they want to relive and enter the episode.

[0229] 2. Submit a request

[0230] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[0231] 3. Receiving and Displaying Content

[0232] The terminal receives the content package sent from the server, checks the received content, plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0233] User Actions

[0234] 1. Select a photo and enter an episode

[0235] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[0236] 2. Emotion analysis

[0237] As users input their stories, the emotion engine recognizes emotions from the input text and the user's facial and vocal expressions. This information is used to customize the memories they relive.

[0238] 3. Re-experiencing memories

[0239] Users can listen to music and smell the scent emitted from the diffuser through their device while viewing photos and detailed story text. For example, while surf rock music plays in the background and the scent of the beach diffuses, they can vividly relive memories of a past beach party by reading the detailed story text. In this case, the content is customized based on emotions, allowing users to relive memories on a deeper level.

[0240] In this way, by combining an emotion engine, the present invention provides a system that provides individual content according to the user's emotional state, allowing the user to vividly relive past experiences. By appealing to multiple senses through music, scent, and text, the entire memory can be reproduced in a richer way, giving the user new emotions.

[0241] The processing flow will be explained below.

[0242] Step 1: Select a photo

[0243] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[0244] Step 2: Episode Input

[0245] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[0246] Step 3: Emotion Recognition

[0247] The device analyzes the episode text entered by the user, as well as the user's facial expressions and voice, and uses an emotion engine to recognize the user's emotions. For example, the emotion "joy" can be recognized.

[0248] Step 4: Create a request

[0249] The terminal packages the selected photo, the input episode, and the recognized emotion information, and creates request data to be sent to the server.

[0250] Step 5: Submitting the request

[0251] The terminal transmits the created request data to the server.

[0252] Step 6: Receiving the request

[0253] The server receives the request data sent from the terminal.

[0254] Step 7: Data analysis

[0255] The server analyzes the received request data, extracting location and time information from the photo metadata, analyzing the episode entered by the user for keywords, and adding any recognized emotional information to the analysis.

[0256] Step 8: Content Generation

[0257] Based on the analyzed data, the server generates music data, scent data, and episode text that match the user's selected photos and episodes, emotions. For example, based on the keywords "beach," "summer," "party," and emotion information such as "joy," a surf rock playlist, beach scent data, and detailed episode text are generated.

[0258] Step 9: Submit content

[0259] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[0260] Step 10: Receive the package

[0261] The terminal receives the content package sent from the server and checks the received content.

[0262] Step 11: Prepare your content

[0263] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0264] Step 12: Emotional Adjustment

[0265] Based on the recognized emotions, the device will appropriately adjust the volume of the music being played, the intensity of the scent, the font size of the episode text being displayed, and so on.

[0266] Step 13: Play and display content

[0267] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[0268] Step 14: Re-experience the memory

[0269] Users relive their memories in a three-dimensional way, customized based on their emotions, through displayed photos, played music, scents, and episode text.

[0270] Example 2

[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0272] In conventional album systems, users' re-experiencing of memories is often limited to visual elements, lacking emotional richness and multi-sensory re-experiencing. Furthermore, it is difficult to generate customized content based on the user's emotions, making it impossible to enhance the vividness and emotion of memories. This has led to issues such as reduced satisfaction when users re-experiencing past memories.

[0273] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0274] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for analyzing emotional information based on the selected image and the input episode and generating associated music data, scent data, and episode text, means for transmitting the generated music data, scent data, and episode text to the user's information processing device, and means for playing the received music data, spraying the scent data with a diffusion device, and displaying the episode text, thereby enabling the user to re-experience past memories more vividly and emotionally through customized multi-sensory content based on emotions.

[0275] "Means for users to select images associated with past memories" refers to tools or functions that allow users to select images associated with specific memories from their own image albums through the system interface.

[0276] The "means for inputting an episode related to an image selected by the user" is a field or interface for inputting in text form the events and emotions that occurred at the time regarding the image selected by the user.

[0277] "Means for analyzing emotional information and generating associated music data, scent data, and episode text" refers to functions and algorithms for analyzing a user's emotions based on input episode text and image metadata, and generating music, scents, and complementary descriptions appropriate to those emotions.

[0278] "Means for transmitting the generated music data, scent data, and episode text to the user's information processing device" refers to a network communication function for transferring various data generated by the server in an appropriate format to the device used by the user.

[0279] "Means for playing received music data, spraying fragrance data by a diffusion device, and displaying episode text" refers to a function that allows the user's device to play the music data received on a music player, spray fragrance based on the fragrance data by a diffusion device, and display the episode text on a display screen.

[0280] "Means for analyzing image metadata and extracting location and time information" refers to functions and algorithms for analyzing metadata such as location information (GPS data) and shooting date and time embedded in image files and extracting the necessary information.

[0281] "Means for analyzing episodes entered by users and generating related detailed information and supplementary episodes" refers to algorithms and functions for analyzing episode text entered by users using natural language processing technology and generating detailed explanations and additional episodes related to the content.

[0282] The present invention relates to an album system that allows users to re-experience past memories in a multi-sensory manner by using a system that combines an emotion analysis engine. The specific operation and program processing of this system are described below.

[0283] Hardware and software used

[0284] 1. Server Hardware:

[0285] Any cloud server (e.g. cloud hosting service)

[0286] 2. Server Software:

[0287] Web server (e.g. Apache, Nginx)

[0288] Sentiment analysis engine (e.g., sentiment analysis API)

[0289] Music generation services (e.g., music streaming APIs)

[0290] Database (e.g. database system)

[0291] 3. Terminal Hardware:

[0292] Smart devices (e.g. smartphones, tablets)

[0293] Bluetooth-enabled diffuser

[0294] 4. Terminal software:

[0295] Mobile applications (e.g., mobile development frameworks)

[0296] Bluetooth API (e.g. Bluetooth communication library)

[0297] Server Processing

[0298] 1. Request acceptance

[0299] The server receives album requests sent by users from their devices. Specifically, it receives HTTP requests and analyzes the image ID, metadata, and episode text entered by the user. For example, it receives a POST request to " / album / request."

[0300] 2. Data Analysis

[0301] The server analyzes the received request data. First, it reads the EXIF ​​data to extract location and time information from the image metadata. Next, it uses a natural language processing (NLP) engine to analyze the episode text entered by the user, and then uses a sentiment analysis engine to identify emotions. For example, the emotion "joy" is recognized from the text "I had a lot of fun."

[0302] 3. Content Generation

[0303] The server generates content based on the analyzed data. It uses a music generation service to obtain a playlist that matches "joy" and "summer." Scent data is obtained using an API that generates beach-related scents. A generative AI model is used to generate detailed episode text. For example, "I enjoyed partying with friends at the beach in the hot sunshine of August 2022."

[0304] 4. Content Submission

[0305] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format, compressed, and transmitted.

[0306] Terminal handling

[0307] 1. Interface provision

[0308] The device provides a UI for users to browse albums and select photos, such as displaying an album list, thumbnails of images, and an episode entry field. This is implemented using a UI framework.

[0309] 2. Submit a request

[0310] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request. Specifically, it sends JSON data containing the selected image ID and episode text in the body of the POST request.

[0311] 3. Receiving and Displaying Content

[0312] The device receives the content package sent from the server, parses the received data in JSON format, plays music using the music playback API, sprays the scent from the diffuser using the Bluetooth API, and displays the episode text in a text view.

[0313] User Actions

[0314] 1. Select an image and enter an episode

[0315] Users select an image related to a past memory from the device's album screen. When they tap the image, a field appears where they can enter the episode text. For example, they could enter, "I had a beach party with my friends. It was so much fun."

[0316] 2. Emotion analysis

[0317] While the user is entering the episode, the sentiment analysis engine automatically analyzes the input text and identifies the sentiment, which is then included in the request sent to the server.

[0318] 3. Re-experiencing memories

[0319] The user relives the memory using the content received on the device. Music plays, the diffuser sprays the scent, and detailed episode text is displayed. For example, listening to surf rock music while feeling the scent of the beach and reading the detailed episode text allows the user to relive the memory more vividly.

[0320] Examples and prompts

[0321] Examples:

[0322] If a user wants to relive a trip to the beach in the summer of 2022, they can do the following:

[0323] 1. The user selects a photo taken at the beach in the summer of 2022 from their album and enters a memory of that time, such as, "I had a great time at a beach party with friends."

[0324] 2. The server receives this information and uses an emotion analysis engine to recognize the emotion "joy."

[0325] 3. Based on the emotion information, the server generates a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded."

[0326] 4. The generated content is sent to the user's device, where the user listens to the music, smells the scent emanating from the diffuser, and enjoys the detailed episode text.

[0327] Prompt statement:

[0328] "Choose a photo you took at the beach in the summer of 2022 and input your memories of having a great time at a beach party with friends. Generate customized music data, scent data, and story text based on the joy analyzed by the emotion engine."

[0329] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0330] Step 1: Request acceptance

[0331] explanation

[0332] The server receives album requests sent from the user's device. Specifically, when the user selects "Photos taken at the beach in summer 2022" and sends a request to recreate that memory, the server receives the request as an HTTP POST request.

[0333] Input and Output

[0334] Input: An HTTP POST request containing the user-selected image ID, metadata, and episode text.

[0335] Output: Analysis result of received request data.

[0336] Specific actions

[0337] The server receives the HTTP POST request and extracts the image ID, metadata, and episode text from the request body.

[0338] Step 2: Analyze the data

[0339] explanation

[0340] The server analyzes the received request data, including reading EXIF ​​data to extract location and time information from the image metadata, then analyzes the episode text entered by the user using a natural language processing (NLP) engine and identifies emotions using a sentiment analysis engine.

[0341] Input and Output

[0342] Input: Image metadata, episode text

[0343] Output: Extracted location information, time information, and identified emotion information

[0344] Specific actions

[0345] The server uses the EXIF ​​library to parse the image metadata and extract location and time information.

[0346] The server analyzes the episode text using an NLP engine and identifies emotional information using a sentiment analysis engine.

[0347] Step 3: Generate content

[0348] explanation

[0349] The server generates content based on the analyzed data. Based on the identified emotional information, it uses a music generation service to obtain an appropriate playlist. Furthermore, it uses an API to generate scent data, and uses a generative AI model to generate detailed episode text.

[0350] Input and Output

[0351] Input: location information, time information, emotion information

[0352] Output: Music data, scent data, episode text

[0353] Specific actions

[0354] The server uses a music streaming API to retrieve a playlist based on emotion information.

[0355] The server uses the scent generation API to obtain beach-related scent data.

[0356] The server uses the generative AI model to generate detailed episode text.

[0357] Step 4: Submit your content

[0358] explanation

[0359] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format and appropriately compressed.

[0360] Input and Output

[0361] Input: Generated music data, scent data, episode text

[0362] Output: Sending the content package to the device

[0363] Specific actions

[0364] The server packages the generated data in JSON format and sends an HTTP response.

[0365] Step 5: Providing an Interface

[0366] explanation

[0367] The device provides an interface for users to browse albums and select images, such as an album list display, image thumbnail display, and episode input field.

[0368] Input and Output

[0369] Input: User actions

[0370] Output: Image selection UI, episode input field

[0371] Specific actions

[0372] The device uses a user interface framework to display image selection and text entry fields.

[0373] Step 6: Submitting the request

[0374] explanation

[0375] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request.

[0376] Input and Output

[0377] Input: User-selected image metadata, episode text

[0378] Output: Sending an HTTP request to the server

[0379] Specific actions

[0380] The device converts the metadata and episode text into JSON format and sends it to the server as an HTTP POST request.

[0381] Step 7: Receiving and displaying content

[0382] explanation

[0383] The device receives the content package sent from the server and displays it in the appropriate format. It plays music using the music playback API and sprays scent from the diffuser using the Bluetooth API. The episode text is displayed in a text view.

[0384] Input and Output

[0385] Input: Received content package (music data, scent data, episode text)

[0386] Output: Playing music, spraying scent, displaying episode text

[0387] Specific actions

[0388] The device parses the received JSON data and plays the music using a music playback library.

[0389] The device uses Bluetooth API to spray the scent from the diffuser.

[0390] The device displays the episode text on the screen.

[0391] (Application example 2)

[0392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0393] When users relive past memories, they need to recreate them more vividly and emotionally through multi-sensory information, including not only visual information but also music and scent. In particular, a system is needed that can analyze the user's emotions and customize content based on that to allow users to relive individual memories more deeply.

[0394] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to select an image related to a past memory, means for the user to input an event related to the selected image, means for generating related music data, scent data, and detailed text based on the selected image and the input event, means for transmitting the generated music data, scent data, and detailed text to the user's information processing device, means for playing the transmitted music data, spraying the scent data using an aroma device, and displaying the detailed text, means for analyzing the emotions of the event input by the user using an emotion analysis engine, and means for customizing the related music data, scent data, and detailed text based on the emotion analysis results. This allows the user to re-experience past memories multisensorily based on emotions.

[0395] A "user" is an individual who uses the system to input images and events in order to relive past memories.

[0396] "Images" are photographs or drawings that users select to associate with past memories.

[0397] An "event" is a past experience or episode that the user enters in relation to the selected image.

[0398] "Music data" refers to music information selected based on the user's emotions and memories.

[0399] "Scent data" is information about scents selected based on the user's emotions and memories, and is emitted through the aroma device.

[0400] "Detailed text" is a detailed description of the memory that is generated based on the events entered by the user.

[0401] "Information processing device" refers to the terminal device used by the user when using this system, such as a smartphone or a personal computer.

[0402] An "aroma device" is a device that sprays a fragrance based on transmitted fragrance data.

[0403] "Sentiment analysis engine" means software or algorithms used to analyze user emotions from user-entered events and other data.

[0404] "Customization" refers to individually adjusting music data, scent data, and detailed text based on the results of sentiment analysis.

[0405] The present invention is a system for allowing users to re-experience past memories in a multi-sensory manner, and provides a method for customizing music data, scent data, and detailed text based on the user's emotions using an emotion analysis engine. Specific embodiments of this system are described in detail below.

[0406] This system consists of a server and the user's information processing device (such as a smartphone). The flow of user operations, data processing on the server, and the actual re-experience is as follows:

[0407] User Actions

[0408] 1. Image selection and event entry

[0409] Using the application on the information processing device, the user selects an image from his or her photo album that relates to a past memory, and enters an event related to the selected image in a text field.

[0410] 2. Submit a request

[0411] The user makes a request to send the image and incident they entered to the server, including the image metadata and the incident they entered.

[0412] Server Processing

[0413] 1. Data Reception and Analysis

[0414] The server analyzes the received request data. First, it analyzes the image metadata to extract location and time information. Second, it uses a sentiment analysis engine to identify the emotion of the event entered by the user.

[0415] The software used is a Python NLP library (e.g., Hugging Face Transformers, Sentiment Analysis model). As a concrete example of sentiment analysis, the following prompt sentence is used:

[0416] I had a great time at a beach party with my friends that day. Analyze the sentiment of this sentence.

[0417] 2. Content Generation

[0418] The server customizes and generates related music data, scent data, and detailed text based on the results of the sentiment analysis. Music data is acquired using the API of a music streaming service (commonly known as a music streaming API), and scent data is generated using an aroma device API. Detailed text is generated based on information related to the user's input.

[0419] Processing of information processing device

[0420] 1. Content Receipt and Preparation

[0421] The information processing device receives the music data, scent data, and detailed text sent from the server. After receiving the data, it plays the music data, sends the scent data to an aroma device (commonly known as a smart diffuser), sprays the scent, and displays the detailed text.

[0422] For example, based on the data generated by the server, the user's smartphone plays music through the speaker and emits the scent of the beach from a synchronized diffuser, allowing the user to vividly relive past memories while reading the detailed text displayed.

[0423] In this way, by incorporating an emotion analysis engine, it is possible to provide customized multi-sensory content that is in line with the user's emotions, allowing them to recreate past memories in a richer way.

[0424] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0425] Step 1:

[0426] A user launches an application on an information processing device, selects an image from an album that relates to a past memory, and then enters an event related to the image in a text field. This input data includes the file path and metadata of the image, as well as the event text entered by the user.

[0427] Step 2:

[0428] The user makes a request to send the selected image and the entered event to the server. This request includes image metadata and the event text. The entered data is sent from the device to the server.

[0429] Step 3:

[0430] The server analyzes the received request data. First, it extracts location and time information from the image metadata. Next, it uses a sentiment analysis engine to analyze the emotion of the event text entered by the user. The input is the submitted request data, and the output is the analyzed emotion (e.g., joy, sadness, etc.).

[0431] Step 4:

[0432] The server generates related music data, scent data, and detailed text based on the results of the sentiment analysis. A music streaming API is used to obtain music data, and a song list is obtained based on keywords and emotions. An aroma device API is used to generate scent data. The detailed text is generated based on the event entered by the user. The input is the sentiment analysis result and the user's event data, and the output is customized music data, scent data, and detailed text.

[0433] Step 5:

[0434] The server packages the generated music data, scent data, and detailed text, and sends them to the user's information processing device. The input is the customized data, and the output is the data to be sent to the terminal.

[0435] Step 6:

[0436] The terminal receives the content package sent from the server. The received content is music data, scent data, and detailed text, and prepares the corresponding action. The input is the data received from the server, and the output is the preparation state for playback.

[0437] Step 7:

[0438] The terminal plays music data, sends scent data to the aroma device to spray the scent, and displays detailed text on the screen. The input is prepared data, and the output is a state in which the user can re-experience past memories.

[0439] In this way, users can vividly relive past memories through emotionally customized multi-sensory content.

[0440] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0442] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0443] [Second embodiment]

[0444] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0445] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0446] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0447] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0448] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0450] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0451] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0452] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0453] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0454] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0455] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0456] The present invention relates to a system that generates music data, scent data, and episode text by having a user select a photo related to a past memory and input an episode related to that photo, and provides these to a terminal operated by the user. The specific operation and program processing of the system are described below.

[0457] Server Processing

[0458] 1. Request acceptance

[0459] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in the summer of 2022" and sends a request to recreate those memories, the server receives the request.

[0460] 2. Data Analysis

[0461] The server analyzes the received request data. At this time, it extracts the metadata of the photo (location information, time information, etc.) and also analyzes the episode entered by the user. For example, it analyzes the episode "I had a beach party with my friends on this day and it was so much fun" and extracts related keywords.

[0462] 3. Content Generation

[0463] Based on the analyzed data, the server generates music data, scent data, and detailed episode text that match the selected photo and episode. For example, based on the keywords "beach," "summer," and "party," the server generates a surf rock playlist that matches a summer beach party, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[0464] 4. Content Submission

[0465] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[0466] Terminal handling

[0467] 1. Interface provision

[0468] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select a photo they want to relive and enter an episode.

[0469] 2. Submit a request

[0470] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[0471] 3. Receiving and Displaying Content

[0472] The device receives the content package sent from the server, checks the received content, plays the received music data, sprays a scent based on the scent data using the diffuser, and displays the episode text on the screen.

[0473] User Actions

[0474] 1. Select a photo and enter an episode

[0475] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[0476] 2. Re-experiencing memories

[0477] Users can enjoy the photos and detailed story text while listening to music and smelling the scent emanating from the diffuser through their device. For example, while reading the detailed story text with surf rock music playing in the background and the scent of the beach diffusing, they can vividly relive memories of past beach parties.

[0478] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive their past memories in a richer way. By combining music, scent, and detailed text, the system recreates the entire memory in three dimensions, providing users with a new experience.

[0479] The processing flow will be explained below.

[0480] Step 1: Select a photo

[0481] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[0482] Step 2: Episode Input

[0483] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[0484] Step 3: Create a request

[0485] The terminal packages the selected photo and the input episode, and creates request data to be sent to the server.

[0486] Step 4: Submitting the request

[0487] The terminal transmits the created request data to the server.

[0488] Step 5: Receiving the request

[0489] The server receives the request data sent from the terminal.

[0490] Step 6: Data analysis

[0491] The server analyzes the received request data, extracts location and time information from the photo metadata, and analyzes the episodes entered by the user for keywords.

[0492] Step 7: Content Generation

[0493] The server generates music data, scent data, and episode text based on the analyzed data. For example, based on keywords such as "beach," "summer," and "party," the server generates a surf rock playlist, beach scent data, and detailed episode text.

[0494] Step 8: Submit content

[0495] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[0496] Step 9: Receive the package

[0497] The terminal receives the content package sent from the server and checks the received content.

[0498] Step 10: Prepare your content

[0499] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0500] Step 11: Play content

[0501] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[0502] Step 12: Re-experience the memory

[0503] Users re-experience past memories in three dimensions through displayed photos, played music, scents, and episode text.

[0504] Example 1

[0505] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0506] Conventional photo album viewing systems only allow users to visually re-experience photos, and lack the means to recreate memories using senses other than sight. In particular, there is a demand for systems that allow users to re-experience past memories in a more three-dimensional way by combining sensory elements such as music and scent. A method to meet this demand is needed.

[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0508] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for generating associated sound data, scent data, and episode text based on the selected image and the input episode, means for transmitting the generated sound data, scent data, and episode text to the user's output device, and means for playing the transmitted sound data, spraying the scent data with a spraying device, and displaying the episode text, thereby enabling the user to re-experience memories not only through images but also through music and scents.

[0509] A "user" is someone who uses the system to re-experience past memories.

[0510] "Images" are media containing visual information that a user selects to associate with past memories.

[0511] An "episode" is descriptive information about memories or experiences that a user enters in relation to a selected image.

[0512] "Sound data" is data that includes information about music or other sounds associated with a user's memories.

[0513] "Scent data" is data that includes information about scents and fragrances associated with a user's memories.

[0514] "Episode text" refers to an episode entered by a user and complementary text information generated based on the episode.

[0515] A "server" is a central processing unit that analyzes information sent by a user, generates the necessary content, and sends it to the user's terminal.

[0516] An "output device" is a device that allows the user to re-experience the generated audio data, scent data, and episode text.

[0517] A "spraying device" is a device for diffusing a scent based on scent data.

[0518] The present invention is a system that generates sound data, scent data, and episode text by allowing a user to select an image related to a past memory and input an episode related to that image, and provides these to an output device operated by the user. This system operates in cooperation with a server, a terminal, and the user.

[0519] Server Operation

[0520] Request reception

[0521] The server uses Apache HTTP server software to accept requests sent from users' devices. Specifically, an endpoint is set up using the Python Flask framework, and the user selects "images taken at the beach in the summer of 2022" and submits the episode to receive the request.

[0522] Data analysis

[0523] The server parses the data received by the Flask application, using the Pandas library to extract image metadata (location, time, etc.), and the Natural Language Toolkit (NLTK) to parse the user's episodes, extracting keywords such as "beach," "summer," and "party."

[0524] Content Generation

[0525] The server generates the necessary content based on the analysis results. It uses the Spotify API to generate audio data, a pre-built scent database to generate scent data, and GPT-3 to generate episode text. For example, it generates a surf rock playlist, beach scents, and detailed episode text that match the themes of "beach," "summer," and "party."

[0526] Content Submission

[0527] The server sends the generated audio data, scent data, and episode text to the user's output device, where the data is compressed in gzip format and returned as an HTTP response.

[0528] Device behavior

[0529] Interface provided

[0530] The device provides a React-based web interface that displays file selection buttons and text input fields so users can select images from the album screen and enter episodes.

[0531] Send request

[0532] The device packages the data entered by the user into JSON format and sends it to the server. The selected image metadata and episode text are combined into a single JSON object and sent as an HTTP POST request to the Flask endpoint.

[0533] Receiving and displaying content

[0534] The device receives the content returned from the server. The received data is displayed on the screen using React, the sound data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed in a text area.

[0535] User Actions

[0536] Selecting photos and entering episodes

[0537] Users can select images related to past memories from the device's album display screen. Specifically, they can select "Images taken at the beach in the summer of 2022" and enter an episode such as "I had a beach party with friends on this day and it was so much fun."

[0538] Re-experiencing memories

[0539] Users re-experience the content through their device. While the audio data is played and the scent is released from the diffuser, they read the detailed episode text, vividly recreating past memories. For example, they can re-read the detailed episode text while surf rock music plays in the background and the scent of the beach fills the air.

[0540] Examples and prompts

[0541] Specific examples

[0542] Example of input photo: "Photo taken at the beach in summer 2022"

[0543] Example episode: "I had a beach party with my friends that day and it was so much fun. We surfed and had a great time."

[0544] Prompt Sentence Examples

[0545] Photo episode generation:

[0546] Generate photo episodes.

[0547] Photo content: Partying with friends on the beach

[0548] Episode summary: I had a beach party with my friends that day and it was so much fun. We surfed and had a great time.

[0549] Music Playlist Generation:

[0550] Create a music playlist that will go well with your fun beach party.

[0551] Keywords: beach, summer, party

[0552] Scent data generation:

[0553] Provide the perfect scent data for your beach memories.

[0554] Keywords: beach, summer, party

[0555] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0556] Step 1: Request acceptance

[0557] The server uses Apache HTTP server software to accept requests from users. The user selects "Photos taken at the beach in the summer of 2022" and submits the request by sending that memory as an episode. This request is sent to the server as JSON-formatted data containing an image file and an episode description. The selected image file and episode description are provided as input, and the request content is retained as output.

[0558] Step 2: Data analysis

[0559] The server uses the Flask framework to analyze the received request data. It extracts image data metadata (e.g., location and time information) using the Pandas library, and analyzes the episode sentences using NLTK. Specifically, it obtains the photo location and date and time from the image's EXIF ​​information, and extracts keywords such as "beach," "summer," and "party" from the episode sentences. The input is the image data and the episode sentences, and the output is the extracted metadata and keywords.

[0560] Step 3: Content generation

[0561] The server generates related audio data, scent data, and episode text based on the parsed metadata and keywords. It uses the Spotify API to generate a playlist based on keywords, selects appropriate scents from the scent database, and generates episode text using GPT-3. For example, for the keywords "beach," "summer," and "party," it generates surf rock music, beach scents, and detailed episode text that match the keywords. The input is the parsed metadata and keywords, and the output is audio data, scent data, and episode text.

[0562] Step 4: Submit content

[0563] The server sends the generated acoustic data, scent data, and episode text to the user's device. The data is compressed in gzip format for efficient transfer and returned as an HTTP response. The input is the data obtained in the content generation step, and the output is the data sent to the user's device in compressed form.

[0564] Step 5: Provide an interface

[0565] The device provides a web interface built using React. It displays a screen where users can select an image related to a past memory and enter an episode about it. For example, when a user selects a photo on the album screen, a thumbnail preview is displayed and a text area for entering an episode is displayed. The input is the user's operation, and the output is the screen display.

[0566] Step 6: Submitting the request

[0567] The device packages the data entered by the user in JSON format and sends it to the server. Specifically, it combines the metadata of the selected image and the episode text into a single JSON object and sends it as an HTTP POST request. The input is the image data and episode text entered by the user, and the output is the request data sent to the server.

[0568] Step 7: Receiving and displaying content

[0569] The device receives the content sent from the server and displays it on the screen using React. The audio data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed on the screen. For example, when you press the play button, surf rock music plays, the beach scent rises from the diffuser, and the detailed episode text is displayed. The input is the data received from the server, and the output is the display of the user interface and playback operation based on that data.

[0570] Step 8: Select photos and enter episodes

[0571] The user selects an image related to a past memory from the album display screen of the device and enters an episode about that moment. For example, the user selects "an image taken at the beach in the summer of 2022" and enters an episode such as "I had a beach party with friends that day and it was a lot of fun." The input is a past image and an episode description, and the output is the request data sent to the server.

[0572] Step 9: Re-experience the memory

[0573] The user re-experiences the content through the device. The audio data sent from the server is played, and the diffuser emits a scent, while the user reads the episode text and relives past memories. For example, the user re-reads a detailed episode while listening to surf rock music and smelling the scent of the beach. The input is the content displayed on the device, and the output is the sensory memory experience that the user re-experiences.

[0574] (Application example 1)

[0575] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0576] Conventional album viewing systems only allowed users to relive past memories using photos, making it difficult to relive memories that involved sensory elements. Furthermore, photos and text alone did not allow users to relive memories through various senses, such as the atmosphere, scent, and music of the place. This limited the user's experience, making it difficult to relive memories in a richer way.

[0577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0578] In this invention, the server includes means for a user to select an image related to a past record, means for the user to input an event related to the selected image, means for generating related music data, fragrance data, and event text based on the selected image and the input event, means for transmitting the generated music data, fragrance data, and event text to the user's terminal, means for playing the transmitted music data, diffusing the fragrance data by an atomizer, and displaying the event text, and means for reproducing sensory elements based on the image and event selected by the user, thereby enabling the user to re-experience past memories along with the music, fragrance, and detailed episode text.

[0579] "Past Records" refers to image and photo files saved by users relating to previous events.

[0580] "Image" means digital photographic or pictorial data stored on a User's device that visually represents a past event.

[0581] "Event" refers to textual information that describes a specific experience or situation related to the image the user selected.

[0582] "Music Data" refers to digital music files that provide an acoustic ambience generated based on selected images and input events.

[0583] "Scent data" refers to digital data containing scent information that is generated based on a selected image and an input event.

[0584] "Event text" refers to text data containing detailed event information entered by the user.

[0585] "Terminal" refers to an electronic device operated by a user, such as a smartphone, tablet, or personal computer.

[0586] "Playing music data" refers to the act of outputting digital music files as sound on the user's device.

[0587] "Spreading fragrance data using an atomizer" refers to the act of spreading a fragrance using a device that physically releases fragrance information as digital data as a scent.

[0588] "Recreating sensory elements" refers to the act of allowing the user to experience the atmosphere of a location through various senses, such as music, scent, and text, based on the selected image and input event.

[0589] This invention relates to a system that allows a user to select an image related to a past record and input an event related to that image, thereby generating music data, aroma data, and event text, and providing these to the user's terminal.

[0590] System program description

[0591] The system operates as follows.

[0592] 1. User Interface

[0593] The device provides an interface for users to select historical images and enter events, including a photo selection screen and a text field for entering events.

[0594] 2. Submit a request

[0595] The device sends the image selected by the user and the event entered to the server. The request data includes the image file and the text entered by the user.

[0596] 3. Data analysis and content generation

[0597] The server analyzes the received request data, extracts image metadata (location information, time information, etc.), and generates related keywords based on the event entered by the user. Based on this, music data, scent data, and event text are generated.

[0598] 4. Content Submission

[0599] The server sends the generated music data, aroma data, and event text to the user's terminal, where the data is appropriately compressed and transmitted.

[0600] 5. Playing and Displaying Content

[0601] The terminal plays the received music data, controls the atomizer based on the scent data, and displays the event text on the screen.

[0602] Hardware and software used

[0603] Hardware:

[0604] Smartphones: Providing the user interface and displaying content

[0605] Smart glasses and head-mounted displays: Providing user interfaces and immersive experiences

[0606] Diffuser: Diffuses fragrance based on aroma data

[0607] software:

[0608] Python: Server-side data analysis and content generation

[0609] Requests module: Data communication between the server and the terminal

[0610] Backend (e.g., Flask or Django): Manages request processing and data generation

[0611] Front-end (e.g. React or Vue.js): Building the user interface

[0612] Specific operation example

[0613] A user opens the app on their smartphone, selects an image from a past family trip, and enters an event such as, "On this day, my family visited a theme park and had a wonderful day." The server receives this request, analyzes the image metadata and the event, and generates related fun theme music, popcorn scent data, and detailed episode text for that time. This data is sent to the user's device, which plays music, emits the popcorn scent from a diffuser, and displays the event text on the screen.

[0614] Example prompt sentence:

[0615] Photo: Family Travel Theme Park.jpg

[0616] Episode: On this day, my family and I went to a theme park and had a really fun day.

[0617] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive past memories in a richer way.

[0618] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0619] Step 1:

[0620] The user opens the app on their smartphone, selects an image related to a past record from the album screen, and enters the event related to the image in the text field. User input: Image file and event text. Output: User-selected image and event text.

[0621] Step 2:

[0622] The terminal packages the image and event entered by the user and sends it to the server. Input: Image and event text selected by the user. Data processing: Packages the image file and text. Output: Request data sent to the server.

[0623] Step 3:

[0624] The server analyzes the received request data, extracts image metadata (location, time, etc.), and generates related keywords from the event text entered by the user. Input: Request data. Data operation: Metadata extraction and keyword generation. Output: Identified metadata and keywords.

[0625] Step 4:

[0626] The server generates related music data, aroma data, and detailed event text based on the identified metadata and keywords. Input: Metadata and keywords. Data operation: Generation of music data, aroma data, and event text. Output: Generated music data, aroma data, and event text.

[0627] Step 5:

[0628] The server appropriately compresses the generated music data, aroma data, and event text and sends them to the user's terminal. Input: Generated music data, aroma data, and event text. Data processing: Data compression. Output: Data package sent to the terminal.

[0629] Step 6:

[0630] The terminal unpacks the data package received from the server, plays the music data, controls the sprayer based on the scent data, and displays the event text on the screen. Input: Data package sent from the server. Data processing: Unpacking the data. Output: Music to be played, scent to be diffused, text to be displayed.

[0631] Step 7:

[0632] The user listens to music provided through the device, smells the scent, and reads the detailed event text displayed on the screen, reliving past memories with all five senses. Input: Played music, diffused scent, displayed text. Output: Re-experiencing the user's memories.

[0633] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0634] This invention relates to an album system that allows users to re-experience past memories in a richer, more sensory way by incorporating an emotion engine. This system recognizes the user's emotions and can customize music data, scent data, and episode text based on those emotions. The specific operation and program processing of this system are described below.

[0635] Server Processing

[0636] 1. Request acceptance

[0637] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in summer 2022" and sends a request to recreate those memories, the server receives the request.

[0638] 2. Data Analysis

[0639] The server analyzes the received request data. During this process, it extracts the photo's metadata (location information, time information, etc.) and analyzes the episode entered by the user. It also uses an emotion engine to recognize the user's emotions and adds this information to the analysis. For example, the emotion "joy" is recognized for the episode "I had a beach party with my friends that day and had a lot of fun."

[0640] 3. Content Generation

[0641] Based on the analyzed data, the server generates music data, scent data, and episode text that match the selected photo and episode. The generated content is then customized based on the emotional information recognized by the emotion engine. For example, keywords such as "beach," "summer," and "party" are combined with the emotional information of "joy" to generate a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[0642] 4. Content Submission

[0643] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[0644] Terminal handling

[0645] 1. Interface provision

[0646] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select the photo they want to relive and enter the episode.

[0647] 2. Submit a request

[0648] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[0649] 3. Receiving and Displaying Content

[0650] The terminal receives the content package sent from the server, checks the received content, plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0651] User Actions

[0652] 1. Select a photo and enter an episode

[0653] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[0654] 2. Emotion analysis

[0655] As users input their stories, the emotion engine recognizes emotions from the input text and the user's facial and vocal expressions. This information is used to customize the memories they relive.

[0656] 3. Re-experiencing memories

[0657] Users can listen to music and smell the scent emitted from the diffuser through their device while viewing photos and detailed story text. For example, while surf rock music plays in the background and the scent of the beach diffuses, they can vividly relive memories of a past beach party by reading the detailed story text. In this case, the content is customized based on emotions, allowing users to relive memories on a deeper level.

[0658] In this way, by combining an emotion engine, the present invention provides a system that provides individual content according to the user's emotional state, allowing the user to vividly relive past experiences. By appealing to multiple senses through music, scent, and text, the entire memory can be reproduced in a richer way, giving the user new emotions.

[0659] The processing flow will be explained below.

[0660] Step 1: Select a photo

[0661] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[0662] Step 2: Episode Input

[0663] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[0664] Step 3: Emotion Recognition

[0665] The device analyzes the episode text entered by the user, as well as the user's facial expressions and voice, and uses an emotion engine to recognize the user's emotions. For example, the emotion "joy" can be recognized.

[0666] Step 4: Create a request

[0667] The terminal packages the selected photo, the input episode, and the recognized emotion information, and creates request data to be sent to the server.

[0668] Step 5: Submitting the request

[0669] The terminal transmits the created request data to the server.

[0670] Step 6: Receiving the request

[0671] The server receives the request data sent from the terminal.

[0672] Step 7: Data analysis

[0673] The server analyzes the received request data, extracting location and time information from the photo metadata, analyzing the episode entered by the user for keywords, and adding any recognized emotional information to the analysis.

[0674] Step 8: Content Generation

[0675] Based on the analyzed data, the server generates music data, scent data, and episode text that match the user's selected photos and episodes, emotions. For example, based on the keywords "beach," "summer," "party," and emotion information such as "joy," a surf rock playlist, beach scent data, and detailed episode text are generated.

[0676] Step 9: Submit content

[0677] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[0678] Step 10: Receive the package

[0679] The terminal receives the content package sent from the server and checks the received content.

[0680] Step 11: Prepare your content

[0681] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0682] Step 12: Emotional Adjustment

[0683] Based on the recognized emotions, the device will appropriately adjust the volume of the music being played, the intensity of the scent, the font size of the episode text being displayed, and so on.

[0684] Step 13: Play and display content

[0685] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[0686] Step 14: Re-experience the memory

[0687] Users relive their memories in a three-dimensional way, customized based on their emotions, through displayed photos, played music, scents, and episode text.

[0688] Example 2

[0689] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0690] In conventional album systems, users' re-experiencing of memories is often limited to visual elements, lacking emotional richness and multi-sensory re-experiencing. Furthermore, it is difficult to generate customized content based on the user's emotions, making it impossible to enhance the vividness and emotion of memories. This has led to issues such as reduced satisfaction when users re-experiencing past memories.

[0691] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0692] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for analyzing emotional information based on the selected image and the input episode and generating associated music data, scent data, and episode text, means for transmitting the generated music data, scent data, and episode text to the user's information processing device, and means for playing the received music data, spraying the scent data with a diffusion device, and displaying the episode text, thereby enabling the user to re-experience past memories more vividly and emotionally through customized multi-sensory content based on emotions.

[0693] "Means for users to select images associated with past memories" refers to tools or functions that allow users to select images associated with specific memories from their own image albums through the system interface.

[0694] The "means for inputting an episode related to an image selected by the user" is a field or interface for inputting in text form the events and emotions that occurred at the time regarding the image selected by the user.

[0695] "Means for analyzing emotional information and generating associated music data, scent data, and episode text" refers to functions and algorithms for analyzing a user's emotions based on input episode text and image metadata, and generating music, scents, and complementary descriptions appropriate to those emotions.

[0696] "Means for transmitting the generated music data, scent data, and episode text to the user's information processing device" refers to a network communication function for transferring various data generated by the server in an appropriate format to the device used by the user.

[0697] "Means for playing received music data, spraying fragrance data by a diffusion device, and displaying episode text" refers to a function that allows the user's device to play the music data received on a music player, spray fragrance based on the fragrance data by a diffusion device, and display the episode text on a display screen.

[0698] "Means for analyzing image metadata and extracting location and time information" refers to functions and algorithms for analyzing metadata such as location information (GPS data) and shooting date and time embedded in image files and extracting the necessary information.

[0699] "Means for analyzing episodes entered by users and generating related detailed information and supplementary episodes" refers to algorithms and functions for analyzing episode text entered by users using natural language processing technology and generating detailed explanations and additional episodes related to the content.

[0700] The present invention relates to an album system that allows users to re-experience past memories in a multi-sensory manner by using a system that combines an emotion analysis engine. The specific operation and program processing of this system are described below.

[0701] Hardware and software used

[0702] 1. Server Hardware:

[0703] Any cloud server (e.g. cloud hosting service)

[0704] 2. Server Software:

[0705] Web server (e.g. Apache, Nginx)

[0706] Sentiment analysis engine (e.g., sentiment analysis API)

[0707] Music generation services (e.g., music streaming APIs)

[0708] Database (e.g. database system)

[0709] 3. Terminal Hardware:

[0710] Smart devices (e.g. smartphones, tablets)

[0711] Bluetooth-enabled diffuser

[0712] 4. Terminal software:

[0713] Mobile applications (e.g., mobile development frameworks)

[0714] Bluetooth API (e.g. Bluetooth communication library)

[0715] Server Processing

[0716] 1. Request acceptance

[0717] The server receives album requests sent by users from their devices. Specifically, it receives HTTP requests and analyzes the image ID, metadata, and episode text entered by the user. For example, it receives a POST request to " / album / request."

[0718] 2. Data Analysis

[0719] The server analyzes the received request data. First, it reads the EXIF ​​data to extract location and time information from the image metadata. Next, it uses a natural language processing (NLP) engine to analyze the episode text entered by the user, and then uses a sentiment analysis engine to identify emotions. For example, the emotion "joy" is recognized from the text "I had a lot of fun."

[0720] 3. Content Generation

[0721] The server generates content based on the analyzed data. It uses a music generation service to obtain a playlist that matches "joy" and "summer." Scent data is obtained using an API that generates beach-related scents. A generative AI model is used to generate detailed episode text. For example, "I enjoyed partying with friends at the beach in the hot sunshine of August 2022."

[0722] 4. Content Submission

[0723] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format, compressed, and transmitted.

[0724] Terminal handling

[0725] 1. Interface provision

[0726] The device provides a UI for users to browse albums and select photos, such as displaying an album list, thumbnails of images, and an episode entry field. This is implemented using a UI framework.

[0727] 2. Submit a request

[0728] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request. Specifically, it sends JSON data containing the selected image ID and episode text in the body of the POST request.

[0729] 3. Receiving and Displaying Content

[0730] The device receives the content package sent from the server, parses the received data in JSON format, plays music using the music playback API, sprays the scent from the diffuser using the Bluetooth API, and displays the episode text in a text view.

[0731] User Actions

[0732] 1. Select an image and enter an episode

[0733] Users select an image related to a past memory from the device's album screen. When they tap the image, a field appears where they can enter the episode text. For example, they could enter, "I had a beach party with my friends. It was so much fun."

[0734] 2. Emotion analysis

[0735] While the user is entering the episode, the sentiment analysis engine automatically analyzes the input text and identifies the sentiment, which is then included in the request sent to the server.

[0736] 3. Re-experiencing memories

[0737] The user relives the memory using the content received on the device. Music plays, the diffuser sprays the scent, and detailed episode text is displayed. For example, listening to surf rock music while feeling the scent of the beach and reading the detailed episode text allows the user to relive the memory more vividly.

[0738] Examples and prompts

[0739] Examples:

[0740] If a user wants to relive a trip to the beach in the summer of 2022, they can do the following:

[0741] 1. The user selects a photo taken at the beach in the summer of 2022 from their album and enters a memory of that time, such as, "I had a great time at a beach party with friends."

[0742] 2. The server receives this information and uses an emotion analysis engine to recognize the emotion "joy."

[0743] 3. Based on the emotion information, the server generates a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded."

[0744] 4. The generated content is sent to the user's device, where the user listens to the music, smells the scent emanating from the diffuser, and enjoys the detailed episode text.

[0745] Prompt statement:

[0746] "Choose a photo you took at the beach in the summer of 2022 and input your memories of having a great time at a beach party with friends. Generate customized music data, scent data, and story text based on the joy analyzed by the emotion engine."

[0747] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0748] Step 1: Request acceptance

[0749] explanation

[0750] The server receives album requests sent from the user's device. Specifically, when the user selects "Photos taken at the beach in summer 2022" and sends a request to recreate that memory, the server receives the request as an HTTP POST request.

[0751] Input and Output

[0752] Input: An HTTP POST request containing the user-selected image ID, metadata, and episode text.

[0753] Output: Analysis result of received request data.

[0754] Specific actions

[0755] The server receives the HTTP POST request and extracts the image ID, metadata, and episode text from the request body.

[0756] Step 2: Analyze the data

[0757] explanation

[0758] The server analyzes the received request data, including reading EXIF ​​data to extract location and time information from the image metadata, then analyzes the episode text entered by the user using a natural language processing (NLP) engine and identifies emotions using a sentiment analysis engine.

[0759] Input and Output

[0760] Input: Image metadata, episode text

[0761] Output: Extracted location information, time information, and identified emotion information

[0762] Specific actions

[0763] The server uses the EXIF ​​library to parse the image metadata and extract location and time information.

[0764] The server analyzes the episode text using an NLP engine and identifies emotional information using a sentiment analysis engine.

[0765] Step 3: Generate content

[0766] explanation

[0767] The server generates content based on the analyzed data. Based on the identified emotional information, it uses a music generation service to obtain an appropriate playlist. Furthermore, it uses an API to generate scent data, and uses a generative AI model to generate detailed episode text.

[0768] Input and Output

[0769] Input: location information, time information, emotion information

[0770] Output: Music data, scent data, episode text

[0771] Specific actions

[0772] The server uses a music streaming API to retrieve a playlist based on emotion information.

[0773] The server uses the scent generation API to obtain beach-related scent data.

[0774] The server uses the generative AI model to generate detailed episode text.

[0775] Step 4: Submit your content

[0776] explanation

[0777] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format and appropriately compressed.

[0778] Input and Output

[0779] Input: Generated music data, scent data, episode text

[0780] Output: Sending the content package to the device

[0781] Specific actions

[0782] The server packages the generated data in JSON format and sends an HTTP response.

[0783] Step 5: Providing an Interface

[0784] explanation

[0785] The device provides an interface for users to browse albums and select images, such as an album list display, image thumbnail display, and episode input field.

[0786] Input and Output

[0787] Input: User actions

[0788] Output: Image selection UI, episode input field

[0789] Specific actions

[0790] The device uses a user interface framework to display image selection and text entry fields.

[0791] Step 6: Submitting the request

[0792] explanation

[0793] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request.

[0794] Input and Output

[0795] Input: User-selected image metadata, episode text

[0796] Output: Sending an HTTP request to the server

[0797] Specific actions

[0798] The device converts the metadata and episode text into JSON format and sends it to the server as an HTTP POST request.

[0799] Step 7: Receiving and displaying content

[0800] explanation

[0801] The device receives the content package sent from the server and displays it in the appropriate format. It plays music using the music playback API and sprays scent from the diffuser using the Bluetooth API. The episode text is displayed in a text view.

[0802] Input and Output

[0803] Input: Received content package (music data, scent data, episode text)

[0804] Output: Playing music, spraying scent, displaying episode text

[0805] Specific actions

[0806] The device parses the received JSON data and plays the music using a music playback library.

[0807] The device uses Bluetooth API to spray the scent from the diffuser.

[0808] The device displays the episode text on the screen.

[0809] (Application example 2)

[0810] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0811] When users relive past memories, they need to recreate them more vividly and emotionally through multi-sensory information, including not only visual information but also music and scent. In particular, a system is needed that can analyze the user's emotions and customize content based on that to allow users to relive individual memories more deeply.

[0812] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to select an image related to a past memory, means for the user to input an event related to the selected image, means for generating related music data, scent data, and detailed text based on the selected image and the input event, means for transmitting the generated music data, scent data, and detailed text to the user's information processing device, means for playing the transmitted music data, spraying the scent data using an aroma device, and displaying the detailed text, means for analyzing the emotions of the event input by the user using an emotion analysis engine, and means for customizing the related music data, scent data, and detailed text based on the emotion analysis results. This allows the user to re-experience past memories multisensorily based on emotions.

[0813] A "user" is an individual who uses the system to input images and events in order to relive past memories.

[0814] "Images" are photographs or drawings that users select to associate with past memories.

[0815] An "event" is a past experience or episode that the user enters in relation to the selected image.

[0816] "Music data" refers to music information selected based on the user's emotions and memories.

[0817] "Scent data" is information about scents selected based on the user's emotions and memories, and is emitted through the aroma device.

[0818] "Detailed text" is a detailed description of the memory that is generated based on the events entered by the user.

[0819] "Information processing device" refers to the terminal device used by the user when using this system, such as a smartphone or a personal computer.

[0820] An "aroma device" is a device that sprays a fragrance based on transmitted fragrance data.

[0821] "Sentiment analysis engine" means software or algorithms used to analyze user emotions from user-entered events and other data.

[0822] "Customization" refers to individually adjusting music data, scent data, and detailed text based on the results of sentiment analysis.

[0823] The present invention is a system for allowing users to re-experience past memories in a multi-sensory manner, and provides a method for customizing music data, scent data, and detailed text based on the user's emotions using an emotion analysis engine. Specific embodiments of this system are described in detail below.

[0824] This system consists of a server and the user's information processing device (such as a smartphone). The flow of user operations, data processing on the server, and the actual re-experience is as follows:

[0825] User Actions

[0826] 1. Image selection and event entry

[0827] Using the application on the information processing device, the user selects an image from his or her photo album that relates to a past memory, and enters an event related to the selected image in a text field.

[0828] 2. Submit a request

[0829] The user makes a request to send the image and incident they entered to the server, including the image metadata and the incident they entered.

[0830] Server Processing

[0831] 1. Data Reception and Analysis

[0832] The server analyzes the received request data. First, it analyzes the image metadata to extract location and time information. Second, it uses a sentiment analysis engine to identify the emotion of the event entered by the user.

[0833] The software used is a Python NLP library (e.g., Hugging Face Transformers, Sentiment Analysis model). As a concrete example of sentiment analysis, the following prompt sentence is used:

[0834] I had a great time at a beach party with my friends that day. Analyze the sentiment of this sentence.

[0835] 2. Content Generation

[0836] The server customizes and generates related music data, scent data, and detailed text based on the results of the sentiment analysis. Music data is acquired using the API of a music streaming service (commonly known as a music streaming API), and scent data is generated using an aroma device API. Detailed text is generated based on information related to the user's input.

[0837] Processing of information processing device

[0838] 1. Content Receipt and Preparation

[0839] The information processing device receives the music data, scent data, and detailed text sent from the server. After receiving the data, it plays the music data, sends the scent data to an aroma device (commonly known as a smart diffuser), sprays the scent, and displays the detailed text.

[0840] For example, based on the data generated by the server, the user's smartphone plays music through the speaker and emits the scent of the beach from a synchronized diffuser, allowing the user to vividly relive past memories while reading the detailed text displayed.

[0841] In this way, by incorporating an emotion analysis engine, it is possible to provide customized multi-sensory content that is in line with the user's emotions, allowing them to recreate past memories in a richer way.

[0842] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0843] Step 1:

[0844] A user launches an application on an information processing device, selects an image from an album that relates to a past memory, and then enters an event related to the image in a text field. This input data includes the file path and metadata of the image, as well as the event text entered by the user.

[0845] Step 2:

[0846] The user makes a request to send the selected image and the entered event to the server. This request includes image metadata and the event text. The entered data is sent from the device to the server.

[0847] Step 3:

[0848] The server analyzes the received request data. First, it extracts location and time information from the image metadata. Next, it uses a sentiment analysis engine to analyze the emotion of the event text entered by the user. The input is the submitted request data, and the output is the analyzed emotion (e.g., joy, sadness, etc.).

[0849] Step 4:

[0850] The server generates related music data, scent data, and detailed text based on the results of the sentiment analysis. A music streaming API is used to obtain music data, and a song list is obtained based on keywords and emotions. An aroma device API is used to generate scent data. The detailed text is generated based on the event entered by the user. The input is the sentiment analysis result and the user's event data, and the output is customized music data, scent data, and detailed text.

[0851] Step 5:

[0852] The server packages the generated music data, scent data, and detailed text, and sends them to the user's information processing device. The input is the customized data, and the output is the data to be sent to the terminal.

[0853] Step 6:

[0854] The terminal receives the content package sent from the server. The received content is music data, scent data, and detailed text, and prepares the corresponding action. The input is the data received from the server, and the output is the preparation state for playback.

[0855] Step 7:

[0856] The terminal plays music data, sends scent data to the aroma device to spray the scent, and displays detailed text on the screen. The input is prepared data, and the output is a state in which the user can re-experience past memories.

[0857] In this way, users can vividly relive past memories through emotionally customized multi-sensory content.

[0858] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0859] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0860] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0861] [Third embodiment]

[0862] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0863] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0864] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0865] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0866] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0867] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0868] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0869] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0870] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0871] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0872] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0873] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0874] The present invention relates to a system that generates music data, scent data, and episode text by having a user select a photo related to a past memory and input an episode related to that photo, and provides these to a terminal operated by the user. The specific operation and program processing of the system are described below.

[0875] Server Processing

[0876] 1. Request acceptance

[0877] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in the summer of 2022" and sends a request to recreate those memories, the server receives the request.

[0878] 2. Data Analysis

[0879] The server analyzes the received request data. At this time, it extracts the metadata of the photo (location information, time information, etc.) and also analyzes the episode entered by the user. For example, it analyzes the episode "I had a beach party with my friends on this day and it was so much fun" and extracts related keywords.

[0880] 3. Content Generation

[0881] Based on the analyzed data, the server generates music data, scent data, and detailed episode text that match the selected photo and episode. For example, based on the keywords "beach," "summer," and "party," the server generates a surf rock playlist that matches a summer beach party, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[0882] 4. Content Submission

[0883] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[0884] Terminal handling

[0885] 1. Interface provision

[0886] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select a photo they want to relive and enter an episode.

[0887] 2. Submit a request

[0888] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[0889] 3. Receiving and Displaying Content

[0890] The device receives the content package sent from the server, checks the received content, plays the received music data, sprays a scent based on the scent data using the diffuser, and displays the episode text on the screen.

[0891] User Actions

[0892] 1. Select a photo and enter an episode

[0893] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[0894] 2. Re-experiencing memories

[0895] Users can enjoy the photos and detailed story text while listening to music and smelling the scent emanating from the diffuser through their device. For example, while reading the detailed story text with surf rock music playing in the background and the scent of the beach diffusing, they can vividly relive memories of past beach parties.

[0896] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive their past memories in a richer way. By combining music, scent, and detailed text, the system recreates the entire memory in three dimensions, providing users with a new experience.

[0897] The processing flow will be explained below.

[0898] Step 1: Select a photo

[0899] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[0900] Step 2: Episode Input

[0901] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[0902] Step 3: Create a request

[0903] The terminal packages the selected photo and the input episode, and creates request data to be sent to the server.

[0904] Step 4: Submitting the request

[0905] The terminal transmits the created request data to the server.

[0906] Step 5: Receiving the request

[0907] The server receives the request data sent from the terminal.

[0908] Step 6: Data analysis

[0909] The server analyzes the received request data, extracts location and time information from the photo metadata, and analyzes the episodes entered by the user for keywords.

[0910] Step 7: Content Generation

[0911] The server generates music data, scent data, and episode text based on the analyzed data. For example, based on keywords such as "beach," "summer," and "party," the server generates a surf rock playlist, beach scent data, and detailed episode text.

[0912] Step 8: Submit content

[0913] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[0914] Step 9: Receive the package

[0915] The terminal receives the content package sent from the server and checks the received content.

[0916] Step 10: Prepare your content

[0917] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[0918] Step 11: Play content

[0919] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[0920] Step 12: Re-experience the memory

[0921] Users re-experience past memories in three dimensions through displayed photos, played music, scents, and episode text.

[0922] Example 1

[0923] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0924] Conventional photo album viewing systems only allow users to visually re-experience photos, and lack the means to recreate memories using senses other than sight. In particular, there is a demand for systems that allow users to re-experience past memories in a more three-dimensional way by combining sensory elements such as music and scent. A method to meet this demand is needed.

[0925] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0926] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for generating associated sound data, scent data, and episode text based on the selected image and the input episode, means for transmitting the generated sound data, scent data, and episode text to the user's output device, and means for playing the transmitted sound data, spraying the scent data with a spraying device, and displaying the episode text, thereby enabling the user to re-experience memories not only through images but also through music and scents.

[0927] A "user" is someone who uses the system to re-experience past memories.

[0928] "Images" are media containing visual information that a user selects to associate with past memories.

[0929] An "episode" is descriptive information about memories or experiences that a user enters in relation to a selected image.

[0930] "Sound data" is data that includes information about music or other sounds associated with a user's memories.

[0931] "Scent data" is data that includes information about scents and fragrances associated with a user's memories.

[0932] "Episode text" refers to an episode entered by a user and complementary text information generated based on the episode.

[0933] A "server" is a central processing unit that analyzes information sent by a user, generates the necessary content, and sends it to the user's terminal.

[0934] An "output device" is a device that allows the user to re-experience the generated audio data, scent data, and episode text.

[0935] A "spraying device" is a device for diffusing a scent based on scent data.

[0936] The present invention is a system that generates sound data, scent data, and episode text by allowing a user to select an image related to a past memory and input an episode related to that image, and provides these to an output device operated by the user. This system operates in cooperation with a server, a terminal, and the user.

[0937] Server Operation

[0938] Request reception

[0939] The server uses Apache HTTP server software to accept requests sent from users' devices. Specifically, an endpoint is set up using the Python Flask framework, and the user selects "images taken at the beach in the summer of 2022" and submits the episode to receive the request.

[0940] Data analysis

[0941] The server parses the data received by the Flask application, using the Pandas library to extract image metadata (location, time, etc.), and the Natural Language Toolkit (NLTK) to parse the user's episodes, extracting keywords such as "beach," "summer," and "party."

[0942] Content Generation

[0943] The server generates the necessary content based on the analysis results. It uses the Spotify API to generate audio data, a pre-built scent database to generate scent data, and GPT-3 to generate episode text. For example, it generates a surf rock playlist, beach scents, and detailed episode text that match the themes of "beach," "summer," and "party."

[0944] Content Submission

[0945] The server sends the generated audio data, scent data, and episode text to the user's output device, where the data is compressed in gzip format and returned as an HTTP response.

[0946] Device behavior

[0947] Interface provided

[0948] The device provides a React-based web interface that displays file selection buttons and text input fields so users can select images from the album screen and enter episodes.

[0949] Send request

[0950] The device packages the data entered by the user into JSON format and sends it to the server. The selected image metadata and episode text are combined into a single JSON object and sent as an HTTP POST request to the Flask endpoint.

[0951] Receiving and displaying content

[0952] The device receives the content returned from the server. The received data is displayed on the screen using React, the sound data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed in a text area.

[0953] User Actions

[0954] Selecting photos and entering episodes

[0955] Users can select images related to past memories from the device's album display screen. Specifically, they can select "Images taken at the beach in the summer of 2022" and enter an episode such as "I had a beach party with friends on this day and it was so much fun."

[0956] Re-experiencing memories

[0957] Users re-experience the content through their device. While the audio data is played and the scent is released from the diffuser, they read the detailed episode text, vividly recreating past memories. For example, they can re-read the detailed episode text while surf rock music plays in the background and the scent of the beach fills the air.

[0958] Examples and prompts

[0959] Specific examples

[0960] Example of input photo: "Photo taken at the beach in summer 2022"

[0961] Example episode: "I had a beach party with my friends that day and it was so much fun. We surfed and had a great time."

[0962] Prompt Sentence Examples

[0963] Photo episode generation:

[0964] Generate photo episodes.

[0965] Photo content: Partying with friends on the beach

[0966] Episode summary: I had a beach party with my friends that day and it was so much fun. We surfed and had a great time.

[0967] Music Playlist Generation:

[0968] Create a music playlist that will go well with your fun beach party.

[0969] Keywords: beach, summer, party

[0970] Scent data generation:

[0971] Provide the perfect scent data for your beach memories.

[0972] Keywords: beach, summer, party

[0973] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0974] Step 1: Request acceptance

[0975] The server uses Apache HTTP server software to accept requests from users. The user selects "Photos taken at the beach in the summer of 2022" and submits the request by sending that memory as an episode. This request is sent to the server as JSON-formatted data containing an image file and an episode description. The selected image file and episode description are provided as input, and the request content is retained as output.

[0976] Step 2: Data analysis

[0977] The server uses the Flask framework to analyze the received request data. It extracts image data metadata (e.g., location and time information) using the Pandas library, and analyzes the episode sentences using NLTK. Specifically, it obtains the photo location and date and time from the image's EXIF ​​information, and extracts keywords such as "beach," "summer," and "party" from the episode sentences. The input is the image data and the episode sentences, and the output is the extracted metadata and keywords.

[0978] Step 3: Content generation

[0979] The server generates related audio data, scent data, and episode text based on the parsed metadata and keywords. It uses the Spotify API to generate a playlist based on keywords, selects appropriate scents from the scent database, and generates episode text using GPT-3. For example, for the keywords "beach," "summer," and "party," it generates surf rock music, beach scents, and detailed episode text that match the keywords. The input is the parsed metadata and keywords, and the output is audio data, scent data, and episode text.

[0980] Step 4: Submit content

[0981] The server sends the generated acoustic data, scent data, and episode text to the user's device. The data is compressed in gzip format for efficient transfer and returned as an HTTP response. The input is the data obtained in the content generation step, and the output is the data sent to the user's device in compressed form.

[0982] Step 5: Provide an interface

[0983] The device provides a web interface built using React. It displays a screen where users can select an image related to a past memory and enter an episode about it. For example, when a user selects a photo on the album screen, a thumbnail preview is displayed and a text area for entering an episode is displayed. The input is the user's operation, and the output is the screen display.

[0984] Step 6: Submitting the request

[0985] The device packages the data entered by the user in JSON format and sends it to the server. Specifically, it combines the metadata of the selected image and the episode text into a single JSON object and sends it as an HTTP POST request. The input is the image data and episode text entered by the user, and the output is the request data sent to the server.

[0986] Step 7: Receiving and displaying content

[0987] The device receives the content sent from the server and displays it on the screen using React. The audio data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed on the screen. For example, when you press the play button, surf rock music plays, the beach scent rises from the diffuser, and the detailed episode text is displayed. The input is the data received from the server, and the output is the display of the user interface and playback operation based on that data.

[0988] Step 8: Select photos and enter episodes

[0989] The user selects an image related to a past memory from the album display screen of the device and enters an episode about that moment. For example, the user selects "an image taken at the beach in the summer of 2022" and enters an episode such as "I had a beach party with friends that day and it was a lot of fun." The input is a past image and an episode description, and the output is the request data sent to the server.

[0990] Step 9: Re-experience the memory

[0991] The user re-experiences the content through the device. The audio data sent from the server is played, and the diffuser emits a scent, while the user reads the episode text and relives past memories. For example, the user re-reads a detailed episode while listening to surf rock music and smelling the scent of the beach. The input is the content displayed on the device, and the output is the sensory memory experience that the user re-experiences.

[0992] (Application example 1)

[0993] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0994] Conventional album viewing systems only allowed users to relive past memories using photos, making it difficult to relive memories that involved sensory elements. Furthermore, photos and text alone did not allow users to relive memories through various senses, such as the atmosphere, scent, and music of the place. This limited the user's experience, making it difficult to relive memories in a richer way.

[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0996] In this invention, the server includes means for a user to select an image related to a past record, means for the user to input an event related to the selected image, means for generating related music data, fragrance data, and event text based on the selected image and the input event, means for transmitting the generated music data, fragrance data, and event text to the user's terminal, means for playing the transmitted music data, diffusing the fragrance data by an atomizer, and displaying the event text, and means for reproducing sensory elements based on the image and event selected by the user, thereby enabling the user to re-experience past memories along with the music, fragrance, and detailed episode text.

[0997] "Past Records" refers to image and photo files saved by users relating to previous events.

[0998] "Image" means digital photographic or pictorial data stored on a User's device that visually represents a past event.

[0999] "Event" refers to textual information that describes a specific experience or situation related to the image the user selected.

[1000] "Music Data" refers to digital music files that provide an acoustic ambience generated based on selected images and input events.

[1001] "Scent data" refers to digital data containing scent information that is generated based on a selected image and an input event.

[1002] "Event text" refers to text data containing detailed event information entered by the user.

[1003] "Terminal" refers to an electronic device operated by a user, such as a smartphone, tablet, or personal computer.

[1004] "Playing music data" refers to the act of outputting digital music files as sound on the user's device.

[1005] "Spreading fragrance data using an atomizer" refers to the act of spreading a fragrance using a device that physically releases fragrance information as digital data as a scent.

[1006] "Recreating sensory elements" refers to the act of allowing the user to experience the atmosphere of a location through various senses, such as music, scent, and text, based on the selected image and input event.

[1007] This invention relates to a system that allows a user to select an image related to a past record and input an event related to that image, thereby generating music data, aroma data, and event text, and providing these to the user's terminal.

[1008] System program description

[1009] The system operates as follows.

[1010] 1. User Interface

[1011] The device provides an interface for users to select historical images and enter events, including a photo selection screen and a text field for entering events.

[1012] 2. Submit a request

[1013] The device sends the image selected by the user and the event entered to the server. The request data includes the image file and the text entered by the user.

[1014] 3. Data analysis and content generation

[1015] The server analyzes the received request data, extracts image metadata (location information, time information, etc.), and generates related keywords based on the event entered by the user. Based on this, music data, scent data, and event text are generated.

[1016] 4. Content Submission

[1017] The server sends the generated music data, aroma data, and event text to the user's terminal, where the data is appropriately compressed and transmitted.

[1018] 5. Playing and Displaying Content

[1019] The terminal plays the received music data, controls the atomizer based on the scent data, and displays the event text on the screen.

[1020] Hardware and software used

[1021] Hardware:

[1022] Smartphones: Providing the user interface and displaying content

[1023] Smart glasses and head-mounted displays: Providing user interfaces and immersive experiences

[1024] Diffuser: Diffuses fragrance based on aroma data

[1025] software:

[1026] Python: Server-side data analysis and content generation

[1027] Requests module: Data communication between the server and the terminal

[1028] Backend (e.g., Flask or Django): Manages request processing and data generation

[1029] Front-end (e.g. React or Vue.js): Building the user interface

[1030] Specific operation example

[1031] A user opens the app on their smartphone, selects an image from a past family trip, and enters an event such as, "On this day, my family visited a theme park and had a wonderful day." The server receives this request, analyzes the image metadata and the event, and generates related fun theme music, popcorn scent data, and detailed episode text for that time. This data is sent to the user's device, which plays music, emits the popcorn scent from a diffuser, and displays the event text on the screen.

[1032] Example prompt sentence:

[1033] Photo: Family Travel Theme Park.jpg

[1034] Episode: On this day, my family and I went to a theme park and had a really fun day.

[1035] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive past memories in a richer way.

[1036] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1037] Step 1:

[1038] The user opens the app on their smartphone, selects an image related to a past record from the album screen, and enters the event related to the image in the text field. User input: Image file and event text. Output: User-selected image and event text.

[1039] Step 2:

[1040] The terminal packages the image and event entered by the user and sends it to the server. Input: Image and event text selected by the user. Data processing: Packages the image file and text. Output: Request data sent to the server.

[1041] Step 3:

[1042] The server analyzes the received request data, extracts image metadata (location, time, etc.), and generates related keywords from the event text entered by the user. Input: Request data. Data operation: Metadata extraction and keyword generation. Output: Identified metadata and keywords.

[1043] Step 4:

[1044] The server generates related music data, aroma data, and detailed event text based on the identified metadata and keywords. Input: Metadata and keywords. Data operation: Generation of music data, aroma data, and event text. Output: Generated music data, aroma data, and event text.

[1045] Step 5:

[1046] The server appropriately compresses the generated music data, aroma data, and event text and sends them to the user's terminal. Input: Generated music data, aroma data, and event text. Data processing: Data compression. Output: Data package sent to the terminal.

[1047] Step 6:

[1048] The terminal unpacks the data package received from the server, plays the music data, controls the sprayer based on the scent data, and displays the event text on the screen. Input: Data package sent from the server. Data processing: Unpacking the data. Output: Music to be played, scent to be diffused, text to be displayed.

[1049] Step 7:

[1050] The user listens to music provided through the device, smells the scent, and reads the detailed event text displayed on the screen, reliving past memories with all five senses. Input: Played music, diffused scent, displayed text. Output: Re-experiencing the user's memories.

[1051] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1052] This invention relates to an album system that allows users to re-experience past memories in a richer, more sensory way by incorporating an emotion engine. This system recognizes the user's emotions and can customize music data, scent data, and episode text based on those emotions. The specific operation and program processing of this system are described below.

[1053] Server Processing

[1054] 1. Request acceptance

[1055] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in summer 2022" and sends a request to recreate those memories, the server receives the request.

[1056] 2. Data Analysis

[1057] The server analyzes the received request data. During this process, it extracts the photo's metadata (location information, time information, etc.) and analyzes the episode entered by the user. It also uses an emotion engine to recognize the user's emotions and adds this information to the analysis. For example, the emotion "joy" is recognized for the episode "I had a beach party with my friends that day and had a lot of fun."

[1058] 3. Content Generation

[1059] Based on the analyzed data, the server generates music data, scent data, and episode text that match the selected photo and episode. The generated content is then customized based on the emotional information recognized by the emotion engine. For example, keywords such as "beach," "summer," and "party" are combined with the emotional information of "joy" to generate a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[1060] 4. Content Submission

[1061] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[1062] Terminal handling

[1063] 1. Interface provision

[1064] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select the photo they want to relive and enter the episode.

[1065] 2. Submit a request

[1066] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[1067] 3. Receiving and Displaying Content

[1068] The terminal receives the content package sent from the server, checks the received content, plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[1069] User Actions

[1070] 1. Select a photo and enter an episode

[1071] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[1072] 2. Emotion analysis

[1073] As users input their stories, the emotion engine recognizes emotions from the input text and the user's facial and vocal expressions. This information is used to customize the memories they relive.

[1074] 3. Re-experiencing memories

[1075] Users can listen to music and smell the scent emitted from the diffuser through their device while viewing photos and detailed story text. For example, while surf rock music plays in the background and the scent of the beach diffuses, they can vividly relive memories of a past beach party by reading the detailed story text. In this case, the content is customized based on emotions, allowing users to relive memories on a deeper level.

[1076] In this way, by combining an emotion engine, the present invention provides a system that provides individual content according to the user's emotional state, allowing the user to vividly relive past experiences. By appealing to multiple senses through music, scent, and text, the entire memory can be reproduced in a richer way, giving the user new emotions.

[1077] The processing flow will be explained below.

[1078] Step 1: Select a photo

[1079] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[1080] Step 2: Episode Input

[1081] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[1082] Step 3: Emotion Recognition

[1083] The device analyzes the episode text entered by the user, as well as the user's facial expressions and voice, and uses an emotion engine to recognize the user's emotions. For example, the emotion "joy" can be recognized.

[1084] Step 4: Create a request

[1085] The terminal packages the selected photo, the input episode, and the recognized emotion information, and creates request data to be sent to the server.

[1086] Step 5: Submitting the request

[1087] The terminal transmits the created request data to the server.

[1088] Step 6: Receiving the request

[1089] The server receives the request data sent from the terminal.

[1090] Step 7: Data analysis

[1091] The server analyzes the received request data, extracting location and time information from the photo metadata, analyzing the episode entered by the user for keywords, and adding any recognized emotional information to the analysis.

[1092] Step 8: Content Generation

[1093] Based on the analyzed data, the server generates music data, scent data, and episode text that match the user's selected photos and episodes, emotions. For example, based on the keywords "beach," "summer," "party," and emotion information such as "joy," a surf rock playlist, beach scent data, and detailed episode text are generated.

[1094] Step 9: Submit content

[1095] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[1096] Step 10: Receive the package

[1097] The terminal receives the content package sent from the server and checks the received content.

[1098] Step 11: Prepare your content

[1099] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[1100] Step 12: Emotional Adjustment

[1101] Based on the recognized emotions, the device will appropriately adjust the volume of the music being played, the intensity of the scent, the font size of the episode text being displayed, and so on.

[1102] Step 13: Play and display content

[1103] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[1104] Step 14: Re-experience the memory

[1105] Users relive their memories in a three-dimensional way, customized based on their emotions, through displayed photos, played music, scents, and episode text.

[1106] Example 2

[1107] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1108] In conventional album systems, users' re-experiencing of memories is often limited to visual elements, lacking emotional richness and multi-sensory re-experiencing. Furthermore, it is difficult to generate customized content based on the user's emotions, making it impossible to enhance the vividness and emotion of memories. This has led to issues such as reduced satisfaction when users re-experiencing past memories.

[1109] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1110] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for analyzing emotional information based on the selected image and the input episode and generating associated music data, scent data, and episode text, means for transmitting the generated music data, scent data, and episode text to the user's information processing device, and means for playing the received music data, spraying the scent data with a diffusion device, and displaying the episode text, thereby enabling the user to re-experience past memories more vividly and emotionally through customized multi-sensory content based on emotions.

[1111] "Means for users to select images associated with past memories" refers to tools or functions that allow users to select images associated with specific memories from their own image albums through the system interface.

[1112] The "means for inputting an episode related to an image selected by the user" is a field or interface for inputting in text form the events and emotions that occurred at the time regarding the image selected by the user.

[1113] "Means for analyzing emotional information and generating associated music data, scent data, and episode text" refers to functions and algorithms for analyzing a user's emotions based on input episode text and image metadata, and generating music, scents, and complementary descriptions appropriate to those emotions.

[1114] "Means for transmitting the generated music data, scent data, and episode text to the user's information processing device" refers to a network communication function for transferring various data generated by the server in an appropriate format to the device used by the user.

[1115] "Means for playing received music data, spraying fragrance data by a diffusion device, and displaying episode text" refers to a function that allows the user's device to play the music data received on a music player, spray fragrance based on the fragrance data by a diffusion device, and display the episode text on a display screen.

[1116] "Means for analyzing image metadata and extracting location and time information" refers to functions and algorithms for analyzing metadata such as location information (GPS data) and shooting date and time embedded in image files and extracting the necessary information.

[1117] "Means for analyzing episodes entered by users and generating related detailed information and supplementary episodes" refers to algorithms and functions for analyzing episode text entered by users using natural language processing technology and generating detailed explanations and additional episodes related to the content.

[1118] The present invention relates to an album system that allows users to re-experience past memories in a multi-sensory manner by using a system that combines an emotion analysis engine. The specific operation and program processing of this system are described below.

[1119] Hardware and software used

[1120] 1. Server Hardware:

[1121] Any cloud server (e.g. cloud hosting service)

[1122] 2. Server Software:

[1123] Web server (e.g. Apache, Nginx)

[1124] Sentiment analysis engine (e.g., sentiment analysis API)

[1125] Music generation services (e.g., music streaming APIs)

[1126] Database (e.g. database system)

[1127] 3. Terminal Hardware:

[1128] Smart devices (e.g. smartphones, tablets)

[1129] Bluetooth-enabled diffuser

[1130] 4. Terminal software:

[1131] Mobile applications (e.g., mobile development frameworks)

[1132] Bluetooth API (e.g. Bluetooth communication library)

[1133] Server Processing

[1134] 1. Request acceptance

[1135] The server receives album requests sent by users from their devices. Specifically, it receives HTTP requests and analyzes the image ID, metadata, and episode text entered by the user. For example, it receives a POST request to " / album / request."

[1136] 2. Data Analysis

[1137] The server analyzes the received request data. First, it reads the EXIF ​​data to extract location and time information from the image metadata. Next, it uses a natural language processing (NLP) engine to analyze the episode text entered by the user, and then uses a sentiment analysis engine to identify emotions. For example, the emotion "joy" is recognized from the text "I had a lot of fun."

[1138] 3. Content Generation

[1139] The server generates content based on the analyzed data. It uses a music generation service to obtain a playlist that matches "joy" and "summer." Scent data is obtained using an API that generates beach-related scents. A generative AI model is used to generate detailed episode text. For example, "I enjoyed partying with friends at the beach in the hot sunshine of August 2022."

[1140] 4. Content Submission

[1141] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format, compressed, and transmitted.

[1142] Terminal handling

[1143] 1. Interface provision

[1144] The device provides a UI for users to browse albums and select photos, such as displaying an album list, thumbnails of images, and an episode entry field. This is implemented using a UI framework.

[1145] 2. Submit a request

[1146] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request. Specifically, it sends JSON data containing the selected image ID and episode text in the body of the POST request.

[1147] 3. Receiving and Displaying Content

[1148] The device receives the content package sent from the server, parses the received data in JSON format, plays music using the music playback API, sprays the scent from the diffuser using the Bluetooth API, and displays the episode text in a text view.

[1149] User Actions

[1150] 1. Select an image and enter an episode

[1151] Users select an image related to a past memory from the device's album screen. When they tap the image, a field appears where they can enter the episode text. For example, they could enter, "I had a beach party with my friends. It was so much fun."

[1152] 2. Emotion analysis

[1153] While the user is entering the episode, the sentiment analysis engine automatically analyzes the input text and identifies the sentiment, which is then included in the request sent to the server.

[1154] 3. Re-experiencing memories

[1155] The user relives the memory using the content received on the device. Music plays, the diffuser sprays the scent, and detailed episode text is displayed. For example, listening to surf rock music while feeling the scent of the beach and reading the detailed episode text allows the user to relive the memory more vividly.

[1156] Examples and prompts

[1157] Examples:

[1158] If a user wants to relive a trip to the beach in the summer of 2022, they can do the following:

[1159] 1. The user selects a photo taken at the beach in the summer of 2022 from their album and enters a memory of that time, such as, "I had a great time at a beach party with friends."

[1160] 2. The server receives this information and uses an emotion analysis engine to recognize the emotion "joy."

[1161] 3. Based on the emotion information, the server generates a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded."

[1162] 4. The generated content is sent to the user's device, where the user listens to the music, smells the scent emanating from the diffuser, and enjoys the detailed episode text.

[1163] Prompt statement:

[1164] "Choose a photo you took at the beach in the summer of 2022 and input your memories of having a great time at a beach party with friends. Generate customized music data, scent data, and story text based on the joy analyzed by the emotion engine."

[1165] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1166] Step 1: Request acceptance

[1167] explanation

[1168] The server receives album requests sent from the user's device. Specifically, when the user selects "Photos taken at the beach in summer 2022" and sends a request to recreate that memory, the server receives the request as an HTTP POST request.

[1169] Input and Output

[1170] Input: An HTTP POST request containing the user-selected image ID, metadata, and episode text.

[1171] Output: Analysis result of received request data.

[1172] Specific actions

[1173] The server receives the HTTP POST request and extracts the image ID, metadata, and episode text from the request body.

[1174] Step 2: Analyze the data

[1175] explanation

[1176] The server analyzes the received request data, including reading EXIF ​​data to extract location and time information from the image metadata, then analyzes the episode text entered by the user using a natural language processing (NLP) engine and identifies emotions using a sentiment analysis engine.

[1177] Input and Output

[1178] Input: Image metadata, episode text

[1179] Output: Extracted location information, time information, and identified emotion information

[1180] Specific actions

[1181] The server uses the EXIF ​​library to parse the image metadata and extract location and time information.

[1182] The server analyzes the episode text using an NLP engine and identifies emotional information using a sentiment analysis engine.

[1183] Step 3: Generate content

[1184] explanation

[1185] The server generates content based on the analyzed data. Based on the identified emotional information, it uses a music generation service to obtain an appropriate playlist. Furthermore, it uses an API to generate scent data, and uses a generative AI model to generate detailed episode text.

[1186] Input and Output

[1187] Input: location information, time information, emotion information

[1188] Output: Music data, scent data, episode text

[1189] Specific actions

[1190] The server uses a music streaming API to retrieve a playlist based on emotion information.

[1191] The server uses the scent generation API to obtain beach-related scent data.

[1192] The server uses the generative AI model to generate detailed episode text.

[1193] Step 4: Submit your content

[1194] explanation

[1195] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format and appropriately compressed.

[1196] Input and Output

[1197] Input: Generated music data, scent data, episode text

[1198] Output: Sending the content package to the device

[1199] Specific actions

[1200] The server packages the generated data in JSON format and sends an HTTP response.

[1201] Step 5: Providing an Interface

[1202] explanation

[1203] The device provides an interface for users to browse albums and select images, such as an album list display, image thumbnail display, and episode input field.

[1204] Input and Output

[1205] Input: User actions

[1206] Output: Image selection UI, episode input field

[1207] Specific actions

[1208] The device uses a user interface framework to display image selection and text entry fields.

[1209] Step 6: Submitting the request

[1210] explanation

[1211] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request.

[1212] Input and Output

[1213] Input: User-selected image metadata, episode text

[1214] Output: Sending an HTTP request to the server

[1215] Specific actions

[1216] The device converts the metadata and episode text into JSON format and sends it to the server as an HTTP POST request.

[1217] Step 7: Receiving and displaying content

[1218] explanation

[1219] The device receives the content package sent from the server and displays it in the appropriate format. It plays music using the music playback API and sprays scent from the diffuser using the Bluetooth API. The episode text is displayed in a text view.

[1220] Input and Output

[1221] Input: Received content package (music data, scent data, episode text)

[1222] Output: Playing music, spraying scent, displaying episode text

[1223] Specific actions

[1224] The device parses the received JSON data and plays the music using a music playback library.

[1225] The device uses Bluetooth API to spray the scent from the diffuser.

[1226] The device displays the episode text on the screen.

[1227] (Application example 2)

[1228] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1229] When users relive past memories, they need to recreate them more vividly and emotionally through multi-sensory information, including not only visual information but also music and scent. In particular, a system is needed that can analyze the user's emotions and customize content based on that to allow users to relive individual memories more deeply.

[1230] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to select an image related to a past memory, means for the user to input an event related to the selected image, means for generating related music data, scent data, and detailed text based on the selected image and the input event, means for transmitting the generated music data, scent data, and detailed text to the user's information processing device, means for playing the transmitted music data, spraying the scent data using an aroma device, and displaying the detailed text, means for analyzing the emotions of the event input by the user using an emotion analysis engine, and means for customizing the related music data, scent data, and detailed text based on the emotion analysis results. This allows the user to re-experience past memories multisensorily based on emotions.

[1231] A "user" is an individual who uses the system to input images and events in order to relive past memories.

[1232] "Images" are photographs or drawings that users select to associate with past memories.

[1233] An "event" is a past experience or episode that the user enters in relation to the selected image.

[1234] "Music data" refers to music information selected based on the user's emotions and memories.

[1235] "Scent data" is information about scents selected based on the user's emotions and memories, and is emitted through the aroma device.

[1236] "Detailed text" is a detailed description of the memory that is generated based on the events entered by the user.

[1237] "Information processing device" refers to the terminal device used by the user when using this system, such as a smartphone or a personal computer.

[1238] An "aroma device" is a device that sprays a fragrance based on transmitted fragrance data.

[1239] "Sentiment analysis engine" means software or algorithms used to analyze user emotions from user-entered events and other data.

[1240] "Customization" refers to individually adjusting music data, scent data, and detailed text based on the results of sentiment analysis.

[1241] The present invention is a system for allowing users to re-experience past memories in a multi-sensory manner, and provides a method for customizing music data, scent data, and detailed text based on the user's emotions using an emotion analysis engine. Specific embodiments of this system are described in detail below.

[1242] This system consists of a server and the user's information processing device (such as a smartphone). The flow of user operations, data processing on the server, and the actual re-experience is as follows:

[1243] User Actions

[1244] 1. Image selection and event entry

[1245] Using the application on the information processing device, the user selects an image from his or her photo album that relates to a past memory, and enters an event related to the selected image in a text field.

[1246] 2. Submit a request

[1247] The user makes a request to send the image and incident they entered to the server, including the image metadata and the incident they entered.

[1248] Server Processing

[1249] 1. Data Reception and Analysis

[1250] The server analyzes the received request data. First, it analyzes the image metadata to extract location and time information. Second, it uses a sentiment analysis engine to identify the emotion of the event entered by the user.

[1251] The software used is a Python NLP library (e.g., Hugging Face Transformers, Sentiment Analysis model). As a concrete example of sentiment analysis, the following prompt sentence is used:

[1252] I had a great time at a beach party with my friends that day. Analyze the sentiment of this sentence.

[1253] 2. Content Generation

[1254] The server customizes and generates related music data, scent data, and detailed text based on the results of the sentiment analysis. Music data is acquired using the API of a music streaming service (commonly known as a music streaming API), and scent data is generated using an aroma device API. Detailed text is generated based on information related to the user's input.

[1255] Processing of information processing device

[1256] 1. Content Receipt and Preparation

[1257] The information processing device receives the music data, scent data, and detailed text sent from the server. After receiving the data, it plays the music data, sends the scent data to an aroma device (commonly known as a smart diffuser), sprays the scent, and displays the detailed text.

[1258] For example, based on the data generated by the server, the user's smartphone plays music through the speaker and emits the scent of the beach from a synchronized diffuser, allowing the user to vividly relive past memories while reading the detailed text displayed.

[1259] In this way, by incorporating an emotion analysis engine, it is possible to provide customized multi-sensory content that is in line with the user's emotions, allowing them to recreate past memories in a richer way.

[1260] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1261] Step 1:

[1262] A user launches an application on an information processing device, selects an image from an album that relates to a past memory, and then enters an event related to the image in a text field. This input data includes the file path and metadata of the image, as well as the event text entered by the user.

[1263] Step 2:

[1264] The user makes a request to send the selected image and the entered event to the server. This request includes image metadata and the event text. The entered data is sent from the device to the server.

[1265] Step 3:

[1266] The server analyzes the received request data. First, it extracts location and time information from the image metadata. Next, it uses a sentiment analysis engine to analyze the emotion of the event text entered by the user. The input is the submitted request data, and the output is the analyzed emotion (e.g., joy, sadness, etc.).

[1267] Step 4:

[1268] The server generates related music data, scent data, and detailed text based on the results of the sentiment analysis. A music streaming API is used to obtain music data, and a song list is obtained based on keywords and emotions. An aroma device API is used to generate scent data. The detailed text is generated based on the event entered by the user. The input is the sentiment analysis result and the user's event data, and the output is customized music data, scent data, and detailed text.

[1269] Step 5:

[1270] The server packages the generated music data, scent data, and detailed text, and sends them to the user's information processing device. The input is the customized data, and the output is the data to be sent to the terminal.

[1271] Step 6:

[1272] The terminal receives the content package sent from the server. The received content is music data, scent data, and detailed text, and prepares the corresponding action. The input is the data received from the server, and the output is the preparation state for playback.

[1273] Step 7:

[1274] The terminal plays music data, sends scent data to the aroma device to spray the scent, and displays detailed text on the screen. The input is prepared data, and the output is a state in which the user can re-experience past memories.

[1275] In this way, users can vividly relive past memories through emotionally customized multi-sensory content.

[1276] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1277] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1278] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1279] [Fourth embodiment]

[1280] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1281] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1282] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1283] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1284] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1285] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1286] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1287] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1288] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1289] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1290] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1291] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1292] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1293] The present invention relates to a system that generates music data, scent data, and episode text by having a user select a photo related to a past memory and input an episode related to that photo, and provides these to a terminal operated by the user. The specific operation and program processing of the system are described below.

[1294] Server Processing

[1295] 1. Request acceptance

[1296] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in the summer of 2022" and sends a request to recreate those memories, the server receives the request.

[1297] 2. Data Analysis

[1298] The server analyzes the received request data. At this time, it extracts the metadata of the photo (location information, time information, etc.) and also analyzes the episode entered by the user. For example, it analyzes the episode "I had a beach party with my friends on this day and it was so much fun" and extracts related keywords.

[1299] 3. Content Generation

[1300] Based on the analyzed data, the server generates music data, scent data, and detailed episode text that match the selected photo and episode. For example, based on the keywords "beach," "summer," and "party," the server generates a surf rock playlist that matches a summer beach party, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[1301] 4. Content Submission

[1302] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[1303] Terminal handling

[1304] 1. Interface provision

[1305] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select a photo they want to relive and enter an episode.

[1306] 2. Submit a request

[1307] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[1308] 3. Receiving and Displaying Content

[1309] The device receives the content package sent from the server, checks the received content, plays the received music data, sprays a scent based on the scent data using the diffuser, and displays the episode text on the screen.

[1310] User Actions

[1311] 1. Select a photo and enter an episode

[1312] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[1313] 2. Re-experiencing memories

[1314] Users can enjoy the photos and detailed story text while listening to music and smelling the scent emanating from the diffuser through their device. For example, while reading the detailed story text with surf rock music playing in the background and the scent of the beach diffusing, they can vividly relive memories of past beach parties.

[1315] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive their past memories in a richer way. By combining music, scent, and detailed text, the system recreates the entire memory in three dimensions, providing users with a new experience.

[1316] The processing flow will be explained below.

[1317] Step 1: Select a photo

[1318] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[1319] Step 2: Episode Input

[1320] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[1321] Step 3: Create a request

[1322] The terminal packages the selected photo and the input episode, and creates request data to be sent to the server.

[1323] Step 4: Submitting the request

[1324] The terminal transmits the created request data to the server.

[1325] Step 5: Receiving the request

[1326] The server receives the request data sent from the terminal.

[1327] Step 6: Data analysis

[1328] The server analyzes the received request data, extracts location and time information from the photo metadata, and analyzes the episodes entered by the user for keywords.

[1329] Step 7: Content Generation

[1330] The server generates music data, scent data, and episode text based on the analyzed data. For example, based on keywords such as "beach," "summer," and "party," the server generates a surf rock playlist, beach scent data, and detailed episode text.

[1331] Step 8: Submit content

[1332] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[1333] Step 9: Receive the package

[1334] The terminal receives the content package sent from the server and checks the received content.

[1335] Step 10: Prepare your content

[1336] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[1337] Step 11: Play content

[1338] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[1339] Step 12: Re-experience the memory

[1340] Users re-experience past memories in three dimensions through displayed photos, played music, scents, and episode text.

[1341] Example 1

[1342] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1343] Conventional photo album viewing systems only allow users to visually re-experience photos, and lack the means to recreate memories using senses other than sight. In particular, there is a demand for systems that allow users to re-experience past memories in a more three-dimensional way by combining sensory elements such as music and scent. A method to meet this demand is needed.

[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1345] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for generating associated sound data, scent data, and episode text based on the selected image and the input episode, means for transmitting the generated sound data, scent data, and episode text to the user's output device, and means for playing the transmitted sound data, spraying the scent data with a spraying device, and displaying the episode text, thereby enabling the user to re-experience memories not only through images but also through music and scents.

[1346] A "user" is someone who uses the system to re-experience past memories.

[1347] "Images" are media containing visual information that a user selects to associate with past memories.

[1348] An "episode" is descriptive information about memories or experiences that a user enters in relation to a selected image.

[1349] "Sound data" is data that includes information about music or other sounds associated with a user's memories.

[1350] "Scent data" is data that includes information about scents and fragrances associated with a user's memories.

[1351] "Episode text" refers to an episode entered by a user and complementary text information generated based on the episode.

[1352] A "server" is a central processing unit that analyzes information sent by a user, generates the necessary content, and sends it to the user's terminal.

[1353] An "output device" is a device that allows the user to re-experience the generated audio data, scent data, and episode text.

[1354] A "spraying device" is a device for diffusing a scent based on scent data.

[1355] The present invention is a system that generates sound data, scent data, and episode text by allowing a user to select an image related to a past memory and input an episode related to that image, and provides these to an output device operated by the user. This system operates in cooperation with a server, a terminal, and the user.

[1356] Server Operation

[1357] Request reception

[1358] The server uses Apache HTTP server software to accept requests sent from users' devices. Specifically, an endpoint is set up using the Python Flask framework, and the user selects "images taken at the beach in the summer of 2022" and submits the episode to receive the request.

[1359] Data analysis

[1360] The server parses the data received by the Flask application, using the Pandas library to extract image metadata (location, time, etc.), and the Natural Language Toolkit (NLTK) to parse the user's episodes, extracting keywords such as "beach," "summer," and "party."

[1361] Content Generation

[1362] The server generates the necessary content based on the analysis results. It uses the Spotify API to generate audio data, a pre-built scent database to generate scent data, and GPT-3 to generate episode text. For example, it generates a surf rock playlist, beach scents, and detailed episode text that match the themes of "beach," "summer," and "party."

[1363] Content Submission

[1364] The server sends the generated audio data, scent data, and episode text to the user's output device, where the data is compressed in gzip format and returned as an HTTP response.

[1365] Device behavior

[1366] Interface provided

[1367] The device provides a React-based web interface that displays file selection buttons and text input fields so users can select images from the album screen and enter episodes.

[1368] Send request

[1369] The device packages the data entered by the user into JSON format and sends it to the server. The selected image metadata and episode text are combined into a single JSON object and sent as an HTTP POST request to the Flask endpoint.

[1370] Receiving and displaying content

[1371] The device receives the content returned from the server. The received data is displayed on the screen using React, the sound data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed in a text area.

[1372] User Actions

[1373] Selecting photos and entering episodes

[1374] Users can select images related to past memories from the device's album display screen. Specifically, they can select "Images taken at the beach in the summer of 2022" and enter an episode such as "I had a beach party with friends on this day and it was so much fun."

[1375] Re-experiencing memories

[1376] Users re-experience the content through their device. While the audio data is played and the scent is released from the diffuser, they read the detailed episode text, vividly recreating past memories. For example, they can re-read the detailed episode text while surf rock music plays in the background and the scent of the beach fills the air.

[1377] Examples and prompts

[1378] Specific examples

[1379] Example of input photo: "Photo taken at the beach in summer 2022"

[1380] Example episode: "I had a beach party with my friends that day and it was so much fun. We surfed and had a great time."

[1381] Prompt Sentence Examples

[1382] Photo episode generation:

[1383] Generate photo episodes.

[1384] Photo content: Partying with friends on the beach

[1385] Episode summary: I had a beach party with my friends that day and it was so much fun. We surfed and had a great time.

[1386] Music Playlist Generation:

[1387] Create a music playlist that will go well with your fun beach party.

[1388] Keywords: beach, summer, party

[1389] Scent data generation:

[1390] Provide the perfect scent data for your beach memories.

[1391] Keywords: beach, summer, party

[1392] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1393] Step 1: Request acceptance

[1394] The server uses Apache HTTP server software to accept requests from users. The user selects "Photos taken at the beach in the summer of 2022" and submits the request by sending that memory as an episode. This request is sent to the server as JSON-formatted data containing an image file and an episode description. The selected image file and episode description are provided as input, and the request content is retained as output.

[1395] Step 2: Data analysis

[1396] The server uses the Flask framework to analyze the received request data. It extracts image data metadata (e.g., location and time information) using the Pandas library, and analyzes the episode sentences using NLTK. Specifically, it obtains the photo location and date and time from the image's EXIF ​​information, and extracts keywords such as "beach," "summer," and "party" from the episode sentences. The input is the image data and the episode sentences, and the output is the extracted metadata and keywords.

[1397] Step 3: Content generation

[1398] The server generates related audio data, scent data, and episode text based on the parsed metadata and keywords. It uses the Spotify API to generate a playlist based on keywords, selects appropriate scents from the scent database, and generates episode text using GPT-3. For example, for the keywords "beach," "summer," and "party," it generates surf rock music, beach scents, and detailed episode text that match the keywords. The input is the parsed metadata and keywords, and the output is audio data, scent data, and episode text.

[1399] Step 4: Submit content

[1400] The server sends the generated acoustic data, scent data, and episode text to the user's device. The data is compressed in gzip format for efficient transfer and returned as an HTTP response. The input is the data obtained in the content generation step, and the output is the data sent to the user's device in compressed form.

[1401] Step 5: Provide an interface

[1402] The device provides a web interface built using React. It displays a screen where users can select an image related to a past memory and enter an episode about it. For example, when a user selects a photo on the album screen, a thumbnail preview is displayed and a text area for entering an episode is displayed. The input is the user's operation, and the output is the screen display.

[1403] Step 6: Submitting the request

[1404] The device packages the data entered by the user in JSON format and sends it to the server. Specifically, it combines the metadata of the selected image and the episode text into a single JSON object and sends it as an HTTP POST request. The input is the image data and episode text entered by the user, and the output is the request data sent to the server.

[1405] Step 7: Receiving and displaying content

[1406] The device receives the content sent from the server and displays it on the screen using React. The audio data is played using an HTML5 audio player, the scent data sends instructions to a Bluetooth diffuser to spray the scent, and the episode text is displayed on the screen. For example, when you press the play button, surf rock music plays, the beach scent rises from the diffuser, and the detailed episode text is displayed. The input is the data received from the server, and the output is the display of the user interface and playback operation based on that data.

[1407] Step 8: Select photos and enter episodes

[1408] The user selects an image related to a past memory from the album display screen of the device and enters an episode about that moment. For example, the user selects "an image taken at the beach in the summer of 2022" and enters an episode such as "I had a beach party with friends that day and it was a lot of fun." The input is a past image and an episode description, and the output is the request data sent to the server.

[1409] Step 9: Re-experience the memory

[1410] The user re-experiences the content through the device. The audio data sent from the server is played, and the diffuser emits a scent, while the user reads the episode text and relives past memories. For example, the user re-reads a detailed episode while listening to surf rock music and smelling the scent of the beach. The input is the content displayed on the device, and the output is the sensory memory experience that the user re-experiences.

[1411] (Application example 1)

[1412] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1413] Conventional album viewing systems only allowed users to relive past memories using photos, making it difficult to relive memories that involved sensory elements. Furthermore, photos and text alone did not allow users to relive memories through various senses, such as the atmosphere, scent, and music of the place. This limited the user's experience, making it difficult to relive memories in a richer way.

[1414] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1415] In this invention, the server includes means for a user to select an image related to a past record, means for the user to input an event related to the selected image, means for generating related music data, fragrance data, and event text based on the selected image and the input event, means for transmitting the generated music data, fragrance data, and event text to the user's terminal, means for playing the transmitted music data, diffusing the fragrance data by an atomizer, and displaying the event text, and means for reproducing sensory elements based on the image and event selected by the user, thereby enabling the user to re-experience past memories along with the music, fragrance, and detailed episode text.

[1416] "Past Records" refers to image and photo files saved by users relating to previous events.

[1417] "Image" means digital photographic or pictorial data stored on a User's device that visually represents a past event.

[1418] "Event" refers to textual information that describes a specific experience or situation related to the image the user selected.

[1419] "Music Data" refers to digital music files that provide an acoustic ambience generated based on selected images and input events.

[1420] "Scent data" refers to digital data containing scent information that is generated based on a selected image and an input event.

[1421] "Event text" refers to text data containing detailed event information entered by the user.

[1422] "Terminal" refers to an electronic device operated by a user, such as a smartphone, tablet, or personal computer.

[1423] "Playing music data" refers to the act of outputting digital music files as sound on the user's device.

[1424] "Spreading fragrance data using an atomizer" refers to the act of spreading a fragrance using a device that physically releases fragrance information as digital data as a scent.

[1425] "Recreating sensory elements" refers to the act of allowing the user to experience the atmosphere of a location through various senses, such as music, scent, and text, based on the selected image and input event.

[1426] This invention relates to a system that allows a user to select an image related to a past record and input an event related to that image, thereby generating music data, aroma data, and event text, and providing these to the user's terminal.

[1427] System program description

[1428] The system operates as follows.

[1429] 1. User Interface

[1430] The device provides an interface for users to select historical images and enter events, including a photo selection screen and a text field for entering events.

[1431] 2. Submit a request

[1432] The device sends the image selected by the user and the event entered to the server. The request data includes the image file and the text entered by the user.

[1433] 3. Data analysis and content generation

[1434] The server analyzes the received request data, extracts image metadata (location information, time information, etc.), and generates related keywords based on the event entered by the user. Based on this, music data, scent data, and event text are generated.

[1435] 4. Content Submission

[1436] The server sends the generated music data, aroma data, and event text to the user's terminal, where the data is appropriately compressed and transmitted.

[1437] 5. Playing and Displaying Content

[1438] The terminal plays the received music data, controls the atomizer based on the scent data, and displays the event text on the screen.

[1439] Hardware and software used

[1440] Hardware:

[1441] Smartphones: Providing the user interface and displaying content

[1442] Smart glasses and head-mounted displays: Providing user interfaces and immersive experiences

[1443] Diffuser: Diffuses fragrance based on aroma data

[1444] software:

[1445] Python: Server-side data analysis and content generation

[1446] Requests module: Data communication between the server and the terminal

[1447] Backend (e.g., Flask or Django): Manages request processing and data generation

[1448] Front-end (e.g. React or Vue.js): Building the user interface

[1449] Specific operation example

[1450] A user opens the app on their smartphone, selects an image from a past family trip, and enters an event such as, "On this day, my family visited a theme park and had a wonderful day." The server receives this request, analyzes the image metadata and the event, and generates related fun theme music, popcorn scent data, and detailed episode text for that time. This data is sent to the user's device, which plays music, emits the popcorn scent from a diffuser, and displays the event text on the screen.

[1451] Example prompt sentence:

[1452] Photo: Family Travel Theme Park.jpg

[1453] Episode: On this day, my family and I went to a theme park and had a really fun day.

[1454] In this way, the present invention adds a sensory element to the traditional album browsing experience, allowing users to relive past memories in a richer way.

[1455] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1456] Step 1:

[1457] The user opens the app on their smartphone, selects an image related to a past record from the album screen, and enters the event related to the image in the text field. User input: Image file and event text. Output: User-selected image and event text.

[1458] Step 2:

[1459] The terminal packages the image and event entered by the user and sends it to the server. Input: Image and event text selected by the user. Data processing: Packages the image file and text. Output: Request data sent to the server.

[1460] Step 3:

[1461] The server analyzes the received request data, extracts image metadata (location, time, etc.), and generates related keywords from the event text entered by the user. Input: Request data. Data operation: Metadata extraction and keyword generation. Output: Identified metadata and keywords.

[1462] Step 4:

[1463] The server generates related music data, aroma data, and detailed event text based on the identified metadata and keywords. Input: Metadata and keywords. Data operation: Generation of music data, aroma data, and event text. Output: Generated music data, aroma data, and event text.

[1464] Step 5:

[1465] The server appropriately compresses the generated music data, aroma data, and event text and sends them to the user's terminal. Input: Generated music data, aroma data, and event text. Data processing: Data compression. Output: Data package sent to the terminal.

[1466] Step 6:

[1467] The terminal unpacks the data package received from the server, plays the music data, controls the sprayer based on the scent data, and displays the event text on the screen. Input: Data package sent from the server. Data processing: Unpacking the data. Output: Music to be played, scent to be diffused, text to be displayed.

[1468] Step 7:

[1469] The user listens to music provided through the device, smells the scent, and reads the detailed event text displayed on the screen, reliving past memories with all five senses. Input: Played music, diffused scent, displayed text. Output: Re-experiencing the user's memories.

[1470] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1471] This invention relates to an album system that allows users to re-experience past memories in a richer, more sensory way by incorporating an emotion engine. This system recognizes the user's emotions and can customize music data, scent data, and episode text based on those emotions. The specific operation and program processing of this system are described below.

[1472] Server Processing

[1473] 1. Request acceptance

[1474] The server receives album requests sent from the user's device. For example, if a user selects "Photos taken at the beach in summer 2022" and sends a request to recreate those memories, the server receives the request.

[1475] 2. Data Analysis

[1476] The server analyzes the received request data. During this process, it extracts the photo's metadata (location information, time information, etc.) and analyzes the episode entered by the user. It also uses an emotion engine to recognize the user's emotions and adds this information to the analysis. For example, the emotion "joy" is recognized for the episode "I had a beach party with my friends that day and had a lot of fun."

[1477] 3. Content Generation

[1478] Based on the analyzed data, the server generates music data, scent data, and episode text that match the selected photo and episode. The generated content is then customized based on the emotional information recognized by the emotion engine. For example, keywords such as "beach," "summer," and "party" are combined with the emotional information of "joy" to generate a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded with people."

[1479] 4. Content Submission

[1480] The server sends the generated music data, scent data, and episode text to the user's device, where the data is appropriately compressed and packaged for smooth transfer.

[1481] Terminal handling

[1482] 1. Interface provision

[1483] The device provides an interface for the user to browse the album and select a photo, for example, the device's album screen displays a text field for the user to select the photo they want to relive and enter the episode.

[1484] 2. Submit a request

[1485] The device packages the data entered by the user and sends it to the server. The request data includes the metadata of the selected photo and the episode entered by the user.

[1486] 3. Receiving and Displaying Content

[1487] The terminal receives the content package sent from the server, checks the received content, plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[1488] User Actions

[1489] 1. Select a photo and enter an episode

[1490] Users can select a photo related to a past memory from the device's display screen. For example, they can select "Photos taken at the beach in the summer of 2022" and enter the memory of that time as an episode.

[1491] 2. Emotion analysis

[1492] As users input their stories, the emotion engine recognizes emotions from the input text and the user's facial and vocal expressions. This information is used to customize the memories they relive.

[1493] 3. Re-experiencing memories

[1494] Users can listen to music and smell the scent emitted from the diffuser through their device while viewing photos and detailed story text. For example, while surf rock music plays in the background and the scent of the beach diffuses, they can vividly relive memories of a past beach party by reading the detailed story text. In this case, the content is customized based on emotions, allowing users to relive memories on a deeper level.

[1495] In this way, by combining an emotion engine, the present invention provides a system that provides individual content according to the user's emotional state, allowing the user to vividly relive past experiences. By appealing to multiple senses through music, scent, and text, the entire memory can be reproduced in a richer way, giving the user new emotions.

[1496] The processing flow will be explained below.

[1497] Step 1: Select a photo

[1498] Users can open the album screen on their device and select the photo they want to relive, for example, "Photos taken at the beach in the summer of 2022."

[1499] Step 2: Episode Input

[1500] The user enters an episode related to the selected photo into a text field on the device, for example, "I had a beach party with my friends that day and it was so much fun."

[1501] Step 3: Emotion Recognition

[1502] The device analyzes the episode text entered by the user, as well as the user's facial expressions and voice, and uses an emotion engine to recognize the user's emotions. For example, the emotion "joy" can be recognized.

[1503] Step 4: Create a request

[1504] The terminal packages the selected photo, the input episode, and the recognized emotion information, and creates request data to be sent to the server.

[1505] Step 5: Submitting the request

[1506] The terminal transmits the created request data to the server.

[1507] Step 6: Receiving the request

[1508] The server receives the request data sent from the terminal.

[1509] Step 7: Data analysis

[1510] The server analyzes the received request data, extracting location and time information from the photo metadata, analyzing the episode entered by the user for keywords, and adding any recognized emotional information to the analysis.

[1511] Step 8: Content Generation

[1512] Based on the analyzed data, the server generates music data, scent data, and episode text that match the user's selected photos and episodes, emotions. For example, based on the keywords "beach," "summer," "party," and emotion information such as "joy," a surf rock playlist, beach scent data, and detailed episode text are generated.

[1513] Step 9: Submit content

[1514] The server assembles the generated music data, scent data, and episode text into a single package and transmits it to the user's terminal.

[1515] Step 10: Receive the package

[1516] The terminal receives the content package sent from the server and checks the received content.

[1517] Step 11: Prepare your content

[1518] The terminal plays the received music data, prepares the diffuser to spray the scent based on the scent data, and prepares to display the episode text.

[1519] Step 12: Emotional Adjustment

[1520] Based on the recognized emotions, the device will appropriately adjust the volume of the music being played, the intensity of the scent, the font size of the episode text being displayed, and so on.

[1521] Step 13: Play and display content

[1522] The device displays photos while playing music in the background, diffusing scents from a diffuser, and displaying episode text on the screen. For example, a photo of a beach in summer 2022 could be displayed, with surf rock music playing and the scent of the beach filling the air, while detailed episode text is displayed.

[1523] Step 14: Re-experience the memory

[1524] Users relive their memories in a three-dimensional way, customized based on their emotions, through displayed photos, played music, scents, and episode text.

[1525] Example 2

[1526] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1527] In conventional album systems, users' re-experiencing of memories is often limited to visual elements, lacking emotional richness and multi-sensory re-experiencing. Furthermore, it is difficult to generate customized content based on the user's emotions, making it impossible to enhance the vividness and emotion of memories. This has led to issues such as reduced satisfaction when users re-experiencing past memories.

[1528] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1529] In this invention, the server includes means for a user to select an image associated with a past memory, means for the user to input an episode associated with the selected image, means for analyzing emotional information based on the selected image and the input episode and generating associated music data, scent data, and episode text, means for transmitting the generated music data, scent data, and episode text to the user's information processing device, and means for playing the received music data, spraying the scent data with a diffusion device, and displaying the episode text, thereby enabling the user to re-experience past memories more vividly and emotionally through customized multi-sensory content based on emotions.

[1530] "Means for users to select images associated with past memories" refers to tools or functions that allow users to select images associated with specific memories from their own image albums through the system interface.

[1531] The "means for inputting an episode related to an image selected by the user" is a field or interface for inputting in text form the events and emotions that occurred at the time regarding the image selected by the user.

[1532] "Means for analyzing emotional information and generating associated music data, scent data, and episode text" refers to functions and algorithms for analyzing a user's emotions based on input episode text and image metadata, and generating music, scents, and complementary descriptions appropriate to those emotions.

[1533] "Means for transmitting the generated music data, scent data, and episode text to the user's information processing device" refers to a network communication function for transferring various data generated by the server in an appropriate format to the device used by the user.

[1534] "Means for playing received music data, spraying fragrance data by a diffusion device, and displaying episode text" refers to a function that allows the user's device to play the music data received on a music player, spray fragrance based on the fragrance data by a diffusion device, and display the episode text on a display screen.

[1535] "Means for analyzing image metadata and extracting location and time information" refers to functions and algorithms for analyzing metadata such as location information (GPS data) and shooting date and time embedded in image files and extracting the necessary information.

[1536] "Means for analyzing episodes entered by users and generating related detailed information and supplementary episodes" refers to algorithms and functions for analyzing episode text entered by users using natural language processing technology and generating detailed explanations and additional episodes related to the content.

[1537] The present invention relates to an album system that allows users to re-experience past memories in a multi-sensory manner by using a system that combines an emotion analysis engine. The specific operation and program processing of this system are described below.

[1538] Hardware and software used

[1539] 1. Server Hardware:

[1540] Any cloud server (e.g. cloud hosting service)

[1541] 2. Server Software:

[1542] Web server (e.g. Apache, Nginx)

[1543] Sentiment analysis engine (e.g., sentiment analysis API)

[1544] Music generation services (e.g., music streaming APIs)

[1545] Database (e.g. database system)

[1546] 3. Terminal Hardware:

[1547] Smart devices (e.g. smartphones, tablets)

[1548] Bluetooth-enabled diffuser

[1549] 4. Terminal software:

[1550] Mobile applications (e.g., mobile development frameworks)

[1551] Bluetooth API (e.g. Bluetooth communication library)

[1552] Server Processing

[1553] 1. Request acceptance

[1554] The server receives album requests sent by users from their devices. Specifically, it receives HTTP requests and analyzes the image ID, metadata, and episode text entered by the user. For example, it receives a POST request to " / album / request."

[1555] 2. Data Analysis

[1556] The server analyzes the received request data. First, it reads the EXIF ​​data to extract location and time information from the image metadata. Next, it uses a natural language processing (NLP) engine to analyze the episode text entered by the user, and then uses a sentiment analysis engine to identify emotions. For example, the emotion "joy" is recognized from the text "I had a lot of fun."

[1557] 3. Content Generation

[1558] The server generates content based on the analyzed data. It uses a music generation service to obtain a playlist that matches "joy" and "summer." Scent data is obtained using an API that generates beach-related scents. A generative AI model is used to generate detailed episode text. For example, "I enjoyed partying with friends at the beach in the hot sunshine of August 2022."

[1559] 4. Content Submission

[1560] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format, compressed, and transmitted.

[1561] Terminal handling

[1562] 1. Interface provision

[1563] The device provides a UI for users to browse albums and select photos, such as displaying an album list, thumbnails of images, and an episode entry field. This is implemented using a UI framework.

[1564] 2. Submit a request

[1565] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request. Specifically, it sends JSON data containing the selected image ID and episode text in the body of the POST request.

[1566] 3. Receiving and Displaying Content

[1567] The device receives the content package sent from the server, parses the received data in JSON format, plays music using the music playback API, sprays the scent from the diffuser using the Bluetooth API, and displays the episode text in a text view.

[1568] User Actions

[1569] 1. Select an image and enter an episode

[1570] Users select an image related to a past memory from the device's album screen. When they tap the image, a field appears where they can enter the episode text. For example, they could enter, "I had a beach party with my friends. It was so much fun."

[1571] 2. Emotion analysis

[1572] While the user is entering the episode, the sentiment analysis engine automatically analyzes the input text and identifies the sentiment, which is then included in the request sent to the server.

[1573] 3. Re-experiencing memories

[1574] The user relives the memory using the content received on the device. Music plays, the diffuser sprays the scent, and detailed episode text is displayed. For example, listening to surf rock music while feeling the scent of the beach and reading the detailed episode text allows the user to relive the memory more vividly.

[1575] Examples and prompts

[1576] Examples:

[1577] If a user wants to relive a trip to the beach in the summer of 2022, they can do the following:

[1578] 1. The user selects a photo taken at the beach in the summer of 2022 from their album and enters a memory of that time, such as, "I had a great time at a beach party with friends."

[1579] 2. The server receives this information and uses an emotion analysis engine to recognize the emotion "joy."

[1580] 3. Based on the emotion information, the server generates a surf rock playlist, beach scent data, and detailed episode text such as "The temperature was hot that day, and the beach was crowded."

[1581] 4. The generated content is sent to the user's device, where the user listens to the music, smells the scent emanating from the diffuser, and enjoys the detailed episode text.

[1582] Prompt statement:

[1583] "Choose a photo you took at the beach in the summer of 2022 and input your memories of having a great time at a beach party with friends. Generate customized music data, scent data, and story text based on the joy analyzed by the emotion engine."

[1584] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1585] Step 1: Request acceptance

[1586] explanation

[1587] The server receives album requests sent from the user's device. Specifically, when the user selects "Photos taken at the beach in summer 2022" and sends a request to recreate that memory, the server receives the request as an HTTP POST request.

[1588] Input and Output

[1589] Input: An HTTP POST request containing the user-selected image ID, metadata, and episode text.

[1590] Output: Analysis result of received request data.

[1591] Specific actions

[1592] The server receives the HTTP POST request and extracts the image ID, metadata, and episode text from the request body.

[1593] Step 2: Analyze the data

[1594] explanation

[1595] The server analyzes the received request data, including reading EXIF ​​data to extract location and time information from the image metadata, then analyzes the episode text entered by the user using a natural language processing (NLP) engine and identifies emotions using a sentiment analysis engine.

[1596] Input and Output

[1597] Input: Image metadata, episode text

[1598] Output: Extracted location information, time information, and identified emotion information

[1599] Specific actions

[1600] The server uses the EXIF ​​library to parse the image metadata and extract location and time information.

[1601] The server analyzes the episode text using an NLP engine and identifies emotional information using a sentiment analysis engine.

[1602] Step 3: Generate content

[1603] explanation

[1604] The server generates content based on the analyzed data. Based on the identified emotional information, it uses a music generation service to obtain an appropriate playlist. Furthermore, it uses an API to generate scent data, and uses a generative AI model to generate detailed episode text.

[1605] Input and Output

[1606] Input: location information, time information, emotion information

[1607] Output: Music data, scent data, episode text

[1608] Specific actions

[1609] The server uses a music streaming API to retrieve a playlist based on emotion information.

[1610] The server uses the scent generation API to obtain beach-related scent data.

[1611] The server uses the generative AI model to generate detailed episode text.

[1612] Step 4: Submit your content

[1613] explanation

[1614] The server packages the generated music data, scent data, and episode text and sends them to the device as an HTTP response. The data is packaged in JSON format and appropriately compressed.

[1615] Input and Output

[1616] Input: Generated music data, scent data, episode text

[1617] Output: Sending the content package to the device

[1618] Specific actions

[1619] The server packages the generated data in JSON format and sends an HTTP response.

[1620] Step 5: Providing an Interface

[1621] explanation

[1622] The device provides an interface for users to browse albums and select images, such as an album list display, image thumbnail display, and episode input field.

[1623] Input and Output

[1624] Input: User actions

[1625] Output: Image selection UI, episode input field

[1626] Specific actions

[1627] The device uses a user interface framework to display image selection and text entry fields.

[1628] Step 6: Submitting the request

[1629] explanation

[1630] The device sends the metadata of the image selected by the user and the entered episode text to the server as an HTTP request.

[1631] Input and Output

[1632] Input: User-selected image metadata, episode text

[1633] Output: Sending an HTTP request to the server

[1634] Specific actions

[1635] The device converts the metadata and episode text into JSON format and sends it to the server as an HTTP POST request.

[1636] Step 7: Receiving and displaying content

[1637] explanation

[1638] The device receives the content package sent from the server and displays it in the appropriate format. It plays music using the music playback API and sprays scent from the diffuser using the Bluetooth API. The episode text is displayed in a text view.

[1639] Input and Output

[1640] Input: Received content package (music data, scent data, episode text)

[1641] Output: Playing music, spraying scent, displaying episode text

[1642] Specific actions

[1643] The device parses the received JSON data and plays the music using a music playback library.

[1644] The device uses Bluetooth API to spray the scent from the diffuser.

[1645] The device displays the episode text on the screen.

[1646] (Application example 2)

[1647] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1648] When users relive past memories, they need to recreate them more vividly and emotionally through multi-sensory information, including not only visual information but also music and scent. In particular, a system is needed that can analyze the user's emotions and customize content based on that to allow users to relive individual memories more deeply.

[1649] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to select an image related to a past memory, means for the user to input an event related to the selected image, means for generating related music data, scent data, and detailed text based on the selected image and the input event, means for transmitting the generated music data, scent data, and detailed text to the user's information processing device, means for playing the transmitted music data, spraying the scent data using an aroma device, and displaying the detailed text, means for analyzing the emotions of the event input by the user using an emotion analysis engine, and means for customizing the related music data, scent data, and detailed text based on the emotion analysis results. This allows the user to re-experience past memories multisensorily based on emotions.

[1650] A "user" is an individual who uses the system to input images and events in order to relive past memories.

[1651] "Images" are photographs or drawings that users select to associate with past memories.

[1652] An "event" is a past experience or episode that the user enters in relation to the selected image.

[1653] "Music data" refers to music information selected based on the user's emotions and memories.

[1654] "Scent data" is information about scents selected based on the user's emotions and memories, and is emitted through the aroma device.

[1655] "Detailed text" is a detailed description of the memory that is generated based on the events entered by the user.

[1656] "Information processing device" refers to the terminal device used by the user when using this system, such as a smartphone or a personal computer.

[1657] An "aroma device" is a device that sprays a fragrance based on transmitted fragrance data.

[1658] "Sentiment analysis engine" means software or algorithms used to analyze user emotions from user-entered events and other data.

[1659] "Customization" refers to individually adjusting music data, scent data, and detailed text based on the results of sentiment analysis.

[1660] The present invention is a system for allowing users to re-experience past memories in a multi-sensory manner, and provides a method for customizing music data, scent data, and detailed text based on the user's emotions using an emotion analysis engine. Specific embodiments of this system are described in detail below.

[1661] This system consists of a server and the user's information processing device (such as a smartphone). The flow of user operations, data processing on the server, and the actual re-experience is as follows:

[1662] User Actions

[1663] 1. Image selection and event entry

[1664] Using the application on the information processing device, the user selects an image from his or her photo album that relates to a past memory, and enters an event related to the selected image in a text field.

[1665] 2. Submit a request

[1666] The user makes a request to send the image and incident they entered to the server, including the image metadata and the incident they entered.

[1667] Server Processing

[1668] 1. Data Reception and Analysis

[1669] The server analyzes the received request data. First, it analyzes the image metadata to extract location and time information. Second, it uses a sentiment analysis engine to identify the emotion of the event entered by the user.

[1670] The software used is a Python NLP library (e.g., Hugging Face Transformers, Sentiment Analysis model). As a concrete example of sentiment analysis, the following prompt sentence is used:

[1671] I had a great time at a beach party with my friends that day. Analyze the sentiment of this sentence.

[1672] 2. Content Generation

[1673] The server customizes and generates related music data, scent data, and detailed text based on the results of the sentiment analysis. Music data is acquired using the API of a music streaming service (commonly known as a music streaming API), and scent data is generated using an aroma device API. Detailed text is generated based on information related to the user's input.

[1674] Processing of information processing device

[1675] 1. Content Receipt and Preparation

[1676] The information processing device receives the music data, scent data, and detailed text sent from the server. After receiving the data, it plays the music data, sends the scent data to an aroma device (commonly known as a smart diffuser), sprays the scent, and displays the detailed text.

[1677] For example, based on the data generated by the server, the user's smartphone plays music through the speaker and emits the scent of the beach from a synchronized diffuser, allowing the user to vividly relive past memories while reading the detailed text displayed.

[1678] In this way, by incorporating an emotion analysis engine, it is possible to provide customized multi-sensory content that is in line with the user's emotions, allowing them to recreate past memories in a richer way.

[1679] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1680] Step 1:

[1681] A user launches an application on an information processing device, selects an image from an album that relates to a past memory, and then enters an event related to the image in a text field. This input data includes the file path and metadata of the image, as well as the event text entered by the user.

[1682] Step 2:

[1683] The user makes a request to send the selected image and the entered event to the server. This request includes image metadata and the event text. The entered data is sent from the device to the server.

[1684] Step 3:

[1685] The server analyzes the received request data. First, it extracts location and time information from the image metadata. Next, it uses a sentiment analysis engine to analyze the emotion of the event text entered by the user. The input is the submitted request data, and the output is the analyzed emotion (e.g., joy, sadness, etc.).

[1686] Step 4:

[1687] The server generates related music data, scent data, and detailed text based on the results of the sentiment analysis. A music streaming API is used to obtain music data, and a song list is obtained based on keywords and emotions. An aroma device API is used to generate scent data. The detailed text is generated based on the event entered by the user. The input is the sentiment analysis result and the user's event data, and the output is customized music data, scent data, and detailed text.

[1688] Step 5:

[1689] The server packages the generated music data, scent data, and detailed text, and sends them to the user's information processing device. The input is the customized data, and the output is the data to be sent to the terminal.

[1690] Step 6:

[1691] The terminal receives the content package sent from the server. The received content is music data, scent data, and detailed text, and prepares the corresponding action. The input is the data received from the server, and the output is the preparation state for playback.

[1692] Step 7:

[1693] The terminal plays music data, sends scent data to the aroma device to spray the scent, and displays detailed text on the screen. The input is prepared data, and the output is a state in which the user can re-experience past memories.

[1694] In this way, users can vividly relive past memories through emotionally customized multi-sensory content.

[1695] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1696] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1697] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1698] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1699] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1700] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1701] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1702] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1703] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1704] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1705] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1706] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1707] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1708] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1709] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1710] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1711] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1712] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1713] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1714] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1715] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1716] The following is further disclosed regarding the above embodiment.

[1717] (Claim 1)

[1718] a means for a user to select a photograph associated with a past memory;

[1719] A means for a user to input an episode related to a selected photo;

[1720] means for generating associated music data, scent data, and episode text based on the selected photo and the input episode;

[1721] a means for transmitting the generated music data, scent data, and episode text to a user's terminal;

[1722] means for playing the transmitted music data, spraying the scent data by a diffuser, and displaying the episode text;

[1723] A system including:

[1724] (Claim 2)

[1725] 10. The system of claim 1, further comprising means for analyzing metadata of the photo to extract location and time information.

[1726] (Claim 3)

[1727] The system according to claim 1, further comprising means for analyzing the episode input by the user and generating related detailed information and complementary episodes.

[1728] "Example 1"

[1729] (Claim 1)

[1730] a means for a user to select an image associated with a past memory;

[1731] a means for a user to input an episode related to an image selected by the user;

[1732] means for generating associated sound data, scent data, and episode text based on the selected image and the input episode;

[1733] means for transmitting the generated acoustic data, scent data and episode text to a user output device;

[1734] means for playing the transmitted sound data, spraying the scent data by a spraying device, and displaying the episode text;

[1735] A system including:

[1736] (Claim 2)

[1737] 10. The system of claim 1, further comprising means for analyzing metadata of the image to extract location information and time information.

[1738] (Claim 3)

[1739] The system according to claim 1, further comprising means for analyzing the episode input by the user and generating related detailed information and complementary episodes.

[1740] "Application Example 1"

[1741] (Claim 1)

[1742] means for a user to select an image associated with a past record;

[1743] means for a user to input an incident associated with a selected image;

[1744] means for generating associated music data, aroma data, and event text based on the selected image and the input event;

[1745] means for transmitting the generated music data, aroma data, and event text to a user's terminal;

[1746] means for playing the transmitted music data, diffusing the aroma data by an atomizer, and displaying the event text;

[1747] means for recreating sensory elements based on user-selected images and events;

[1748] A system including:

[1749] (Claim 2)

[1750] 10. The system of claim 1, further comprising means for analyzing image metadata to extract location and time information.

[1751] (Claim 3)

[1752] 10. The system of claim 1, further comprising means for analyzing a user input event and generating related detailed information and complementary events.

[1753] "Example 2: Combining Emotion Engines"

[1754] (Claim 1)

[1755] a means for a user to select an image associated with a past memory;

[1756] a means for a user to input an episode related to an image selected by the user;

[1757] means for analyzing emotion information based on the selected image and the input episode, and generating associated music data, scent data, and episode text;

[1758] means for transmitting the generated music data, scent data, and episode text to a user's information processing device;

[1759] means for playing the received music data, spraying the scent data by a diffuser, and displaying the episode text;

[1760] A system including:

[1761] (Claim 2)

[1762] 10. The system of claim 1, further comprising means for analyzing image metadata to extract location and time information.

[1763] (Claim 3)

[1764] 10. The system of claim 1, further comprising means for analyzing the user's input episode and generating related detailed information and complementary episodes.

[1765] "Application example 2 when combining emotion engines"

[1766] (Claim 1)

[1767] a means for a user to select an image associated with a past memory;

[1768] means for a user to input an incident associated with a selected image;

[1769] means for generating related music data, scent data, and detailed text based on the selected image and the input event;

[1770] means for transmitting the generated music data, scent data and detailed text to a user's information processing device;

[1771] means for playing the transmitted music data, spraying the scent data by an aroma device, and displaying detailed text;

[1772] A means for analyzing the emotions of events entered by a user using a sentiment analysis engine;

[1773] A means for customizing related music data, scent data, and detailed text based on the sentiment analysis results;

[1774] A system including:

[1775] (Claim 2)

[1776] 10. The system of claim 1, further comprising means for analyzing image metadata to extract location and time information.

[1777] (Claim 3)

[1778] The system according to claim 1, further comprising means for analyzing an incident input by a user and generating related detailed information and complementary episodes. [Explanation of symbols]

[1779] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to select a photograph associated with a past memory; A means for a user to input an episode related to a selected photo; means for generating associated music data, scent data, and episode text based on the selected photo and the input episode; a means for transmitting the generated music data, scent data, and episode text to a user's terminal; means for playing the transmitted music data, spraying the scent data by a diffuser, and displaying the episode text; A system including:

2. The system of claim 1 further comprising means for analyzing metadata of the photograph to extract location and time information.

3. The system according to claim 1 , further comprising means for analyzing an episode input by a user and generating related detailed information and complementary episodes.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A