System

The system addresses the challenges of managing and interacting with large photo collections by automatically processing and displaying memorable photos and engaging in real-time conversations with an AI model resembling the user, improving photo management and user experience.

JP2026034097APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137218
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Managing and using large numbers of photos taken with mobile devices is time-consuming, and real-time dialogue systems using deep learning technology are not well-established for daily use, making it difficult to efficiently search and display past photos and engage in natural conversations.

Method used

A system that receives photos from a user's device, automatically processes and composites them using AI algorithms, searches for memorable photos based on date, and generates an AI model resembling the user for real-time conversation.

Benefits of technology

Enables efficient photo management, automatic editing, display of memorable photos, and real-time interaction with an AI model that resembles the user, enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034097000001_ABST
    Figure 2026034097000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system includes a means for receiving a photograph, a means for automatically processing and synthesizing the received photograph, a means for specifying and displaying a memorable photograph related to the day, and a means for generating AI on the basis of the photograph and voice of a user to perform conversation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Nowadays, many users take and share photos using mobile devices, but managing and using these photos poses many challenges. Specifically, manually processing and compositing large numbers of photos is time-consuming, and efficiently searching and displaying past photos is difficult. Furthermore, real-time dialogue systems using deep learning technology are still in their infancy, and concrete methods for using them in daily life have yet to be established. The purpose of this invention is to solve these challenges and provide a new system that enriches users' photo and dialogue experiences. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means. First, a means for receiving photos from a user's device is provided, allowing the user to easily send photos to the system. Next, a means for automatically processing and compositing the received photos is provided, adjusting brightness, contrast, and color tone and applying filters using AI algorithms. Furthermore, a means for searching a database for photos taken on the same date in the past based on the current date and displaying memorable photos to the user is provided. The system also includes a means for using DeepFake technology to generate an AI model that looks exactly like the user based on the user's photos and voice, and for using the AI ​​to engage in natural conversations. This allows users to automatically process photos they take to a high quality, efficiently display memorable photos from the past, and interact with an AI that resembles them in real time.

[0006] "Means for receiving photos" refers to a mechanism for sending and receiving photo files and their metadata from a user's device.

[0007] "Means for automatically processing and compositing received photographs" means devices or software that use artificial intelligence (AI) algorithms to adjust the brightness, contrast, and color of photographs, and apply filters and composite images as needed.

[0008] The "means for identifying and displaying photos with memories related to that day" is a mechanism for searching for photos taken on the same day in the past based on the current date and displaying them to the user.

[0009] "Means for generating AI based on a user's photo and voice and engaging in conversation" refers to a device or software that uses Deep Fake technology to generate an AI model that looks exactly like the user using the user's photo and voice data, and then uses that AI model to engage in natural conversation with the user.

[0010] "Metadata" refers to additional information related to a photo, specifically data such as the date and time the photo was taken, the location where it was taken, and camera settings.

[0011] "Artificial intelligence (AI) algorithms" are software agents that perform specific tasks by processing and analyzing large amounts of data.

[0012] "Deep Fake technology" is a technology that uses machine learning and deep learning to generate and convert images and audio.

[0013] A "user-like AI model" is a digital personality created from a user's photo and voice data, and is designed to look and speak like the user. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0036] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0037] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0038] In addition, the server identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. For example, if a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0039] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, if the user provides their own photo and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to enjoy a conversation with the AI.

[0040] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them using Deep Fake technology, improving the efficiency of photo management and use and enhancing the user experience.

[0041] The processing flow will be explained below.

[0042] Step 1:

[0043] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0044] Step 2:

[0045] The device sends the selected photos and metadata to the server, where the actual uploading process takes place using Wi-Fi or mobile data.

[0046] Step 3:

[0047] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0048] Step 4:

[0049] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0050] Step 5:

[0051] The server stores the AI-processed photos in a final storage area, along with the original metadata, making future searches and display easier.

[0052] Step 6:

[0053] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0054] Step 7:

[0055] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0056] Step 8:

[0057] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0058] Step 9:

[0059] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0060] Step 10:

[0061] The user initiates a dialogue with the AI ​​through the device, and the server receives input from the user in real time and generates a response using the AI ​​model. For example, if you ask, "What's the weather like today?", the AI ​​will provide weather information in the format desired by the user. This dialogue can be done via text or voice, improving the user experience.

[0062] Through these steps, the system enables automatic photo editing, identification and display of memorable photos, and real-time interaction using Deep Fake technology.

[0063] Example 1

[0064] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0065] In recent years, users have been taking and storing a large number of digital photos, which requires time and effort to organize and edit them. Furthermore, it is difficult to easily look back on past photos of memories, creating a demand for new ways of communicating using photos. Furthermore, while there is hope for the generation and use of AI models that can converse in real time using the user's own photos and voice, achieving this requires advanced technology.

[0066] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0067] In this invention, the server includes means for taking and sending photos from a terminal, means for receiving and storing the photos and their metadata on the server, means for automatically processing and synthesizing the received photos using an AI algorithm, means for searching and displaying past memorable photos related to that day from a database, and means for generating an AI model using DeepFake technology based on the user's photos and voice and conducting real-time conversations with the user. This allows users to automate the organization and processing of photos, easily look back on memorable photos from special days, and enjoy a new communication experience of interacting in real time with an AI model that resembles them.

[0068] A "terminal" is an electronic device that a user uses to take a photo and transmit the captured photo data and its metadata to a server.

[0069] "Photo" refers to image data that is taken by a user using a terminal and sent to a server.

[0070] "Metadata" is auxiliary information that accompanies photo data, such as the date and time the photo was taken and the location where it was taken.

[0071] The "server" is a system that receives photo data and metadata sent from the device, temporarily stores them, and then processes and synthesizes this data using AI algorithms.

[0072] An "AI algorithm" is an artificial intelligence calculation method used by the server to automatically process and synthesize photo data.

[0073] "Automatic processing" is a process that uses AI algorithms to automatically edit received photo data, such as adjusting brightness, contrast, and applying filters.

[0074] "Synthesis" is the process of combining multiple photographic data to generate a new image.

[0075] "Memorable photos" are specific photos taken by the user in the past that are related to that day.

[0076] "Deep Fake technology" is a technology that generates a realistic AI model that resembles a user based on the user's photo and voice data.

[0077] An "AI model" is a digital agent that resembles a user and is generated using Deep Fake technology based on the user's photo and voice data.

[0078] "Real-time conversation" is a form of communication in which a user interacts with a generated AI model and receives an immediate response.

[0079] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0080] First, a user takes a photo using a device and sends it to a server. A dedicated application is installed on the device, and the user uploads the photo using this application. The device provides a transmission function that includes photo data and its metadata (date and time of photo, location, etc.). A specific example is when a user takes a photo while traveling and sends it to a server via a dedicated application. For example, a user takes a photo of the Eiffel Tower and sends it to a server using the application.

[0081] The server then temporarily stores the received photos and their metadata. The server then automatically processes and combines the stored data using AI algorithms. Specifically, it uses Python scripts and image processing libraries (e.g., OpenCV) to adjust brightness and contrast, apply filters, or combine multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0082] In addition, the server searches the database for past memory photos related to that day and displays them to the user. This function identifies photos taken on the same day in the past and notifies the user. For example, when a user has a birthday, the server finds photos taken on past birthdays and displays them as "past memories" with a special message.

[0083] Finally, an AI model is generated using DeepFake technology based on the user's photo and voice data, allowing the user to enjoy real-time conversations with a digital agent that resembles them. For example, if a user provides their own face and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to have natural conversations with this AI model.

[0084] An example prompt might be:

[0085] "I want to edit a photo of the Eiffel Tower I took during my trip to make it brighter and with more contrast."

[0086] "Find photos from past birthdays and display them with special messages."

[0087] "I would like to use the photos and voice data I provide to generate an AI model that resembles me and then have a conversation with that AI model."

[0088] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them in real time using Deep Fake technology, which is expected to improve the efficiency of photo organization and editing and enhance the user experience.

[0089] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0090] Step 1:

[0091] The user takes a photo on their device, opens a dedicated application, selects the photo, and presses the send button.

[0092] Input: The photo taken and its metadata (date, timestamp, location, etc.).

[0093] Processing: The device uploads the photo data and metadata to the server via a dedicated application.

[0094] Output: Notification of completion of transmission to the server.

[0095] Specific operation: The user takes a photo of a landscape using the camera app on their smartphone, then opens the dedicated app, taps the "Upload Photo" button, and sends the selected photo using the "Send" button.

[0096] Step 2:

[0097] The server receives the photo data and metadata sent from the device and temporarily stores them.

[0098] Input: Photos and metadata sent from the device.

[0099] Processing: The server receives the HTTP request, analyzes the photo data and metadata, and stores the data in a storage service.

[0100] Output: Notification that the photo data and metadata have been saved to storage.

[0101] Specific operation: The server saves the received photos to a storage service (e.g. Amazon S3). After saving is complete, add a record of the save to the database.

[0102] Step 3:

[0103] The server automatically processes and synthesizes the stored photo data using AI algorithms.

[0104] Input: Saved photo data.

[0105] Processing: The server runs a Python script that uses image processing libraries (e.g. OpenCV) to adjust brightness and contrast, apply filters, and possibly blend multiple photos together.

[0106] Output: Processed photo data.

[0107] What it does: The server processes the stored photos sequentially, applying a sepia filter to a photo of the Eiffel Tower, for example, and adjusting the brightness and contrast of another photo.

[0108] Step 4:

[0109] The server searches the database for past memorable photos related to that day, identifies them, and notifies the user.

[0110] Input: Date and time information for that day.

[0111] Processing: The server executes an SQL query to search the database for previous photos taken on the same day. If any are found, the user is notified.

[0112] Output: Identified past memory photos and notification messages.

[0113] What it does: On the user's birthday, the server finds photos taken on past birthdays and notifies the user via the app with a special message.

[0114] Step 5:

[0115] The user sends their own photo and voice data to the server.

[0116] Input: User photo and voice data.

[0117] Processing: The device uploads the photo and audio data to the server via a dedicated application.

[0118] Output: Notification of completion of transmission to the server.

[0119] Specific operation: The user takes a photo of their face and records their voice using a dedicated app, and then sends the data to the server.

[0120] Step 6:

[0121] Based on the received data, the server uses Deep Fake technology to generate an AI model that resembles the user.

[0122] Input: User photo and voice data.

[0123] Processing: The server uses a DeepFake generation tool (e.g., DeepFaceLab) to generate an AI model, which is synthesized based on the user's photo and voice.

[0124] Output: The generated AI model.

[0125] How it works: The server uses deep learning to generate a realistic AI model based on the user's photo and voice data. This AI model is stored on the server and used to interact with the user.

[0126] Step 7:

[0127] Users interact with the generated AI model in real time.

[0128] Input: The generated AI model.

[0129] Processing: The user initiates a conversation with the AI ​​model using a dedicated app. The server runs the necessary infrastructure (e.g., a chat engine) and responds immediately to the user's questions.

[0130] Output: Interaction logs and responses with the AI ​​model.

[0131] How it works: Users communicate with the AI ​​model through text and voice chat. For example, they can ask, "What's the weather like today?" and the AI ​​model will respond in real time.

[0132] Through these steps, users can automatically enhance their photos, easily relive memorable photos from special occasions, and even interact with an AI model that resembles them in real time.

[0133] (Application example 1)

[0134] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0135] Conventional photo editing systems and conversational AI systems have limited advanced interaction and utilization using user photos and voice data, particularly in the virtual store try-on experience. The present invention aims to solve these problems by providing a system that uses user photos and voice data to achieve more advanced photo editing and natural dialogue with a digital agent, and further improves the virtual store try-on experience.

[0136] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0137] In this invention, the server includes means for receiving photos, means for automatically processing and combining the received photos, means for identifying and displaying photos that are memorable for that day, means for generating an AI based on the user's photos and voice to hold a conversation, and means for generating a virtual avatar based on the user's photos and voice to assist in trying on products in a virtual store. This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home, allowing them to check the look and fit of products.

[0138] "Means for receiving photos" means any device or software for sending and receiving photos and their metadata from a user's terminal to the server.

[0139] "Means for automatically processing and synthesizing received photographs" refers to devices or software that use AI algorithms on received photographic data to automatically adjust brightness, contrast, color tone, and apply filters.

[0140] "Means for identifying and displaying memorable photographs relating to that date" refers to a device or software for searching a database for past photographs relating to a particular date and displaying them to a user.

[0141] "Means for generating AI based on a user's photo and voice to converse" refers to a device or software that uses photo and voice data obtained from a user to generate a digital agent that resembles the user using Deep Fake technology, allowing the user to converse naturally with that agent.

[0142] "Means for generating a virtual avatar and assisting in trying on products in a virtual store" refers to a device or software that generates a digital avatar based on a user's photograph and voice, and enables the user to simulate trying on products in a virtual store using that avatar.

[0143] This invention provides an advanced photo processing and AI dialogue system that utilizes photos and voice data taken by users. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0144] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0145] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms. This includes adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by the user and adjusts color and contrast to create a clearer image. The server also identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on that day and displays them with a special message.

[0146] Another distinctive feature of this invention is the ability to use DeepFake technology to generate a digital avatar based on the user's photo and voice data, helping them try on products in a virtual store. Users take photos of themselves with their smartphones and provide voice data. This generates a digital avatar, allowing them to try on products in the virtual store. The generated digital avatar allows the user to try on selected products in real time and check their appearance and fit.

[0147] Hardware and software used:

[0148] Smartphone or PC (used for capturing photos and audio)

[0149] Camera (takes photos)

[0150] Microphone (acquires audio data)

[0151] Server (temporary storage of photo and audio data, processing, generation of virtual avatars)

[0152] OpenCV (used for capturing and displaying photos)

[0153] Deep Fake technology (used to generate digital avatars)

[0154] Python script (used to coordinate the entire process)

[0155] Examples:

[0156] Let's say a user wants to try on new clothes. They launch the app and capture their photo and audio data. Based on this, the app generates a digital avatar and shows them how the selected clothes will look in real time. Through virtual try-on, users can see how the products will look at home without having to go to a physical store.

[0157] Example prompt sentence:

[0158] Prompt: "Generate a digital agent that resembles the user based on the following photo and audio data. The photo and audio data are attached. This agent will act as an avatar for the user to try on products in a virtual store."

[0159] This approach allows users to enjoy a virtual store experience and avoids the hassle of visiting a physical store.

[0160] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0161] Step 1:

[0162] Photo and audio data capture

[0163] Users can take their own photos and record audio data using a smartphone or PC, using the camera to capture the photos and the microphone to capture the audio data.

[0164] Input: Smartphone, PC, camera, microphone

[0165] Output: Photo data, audio data

[0166] Step 2:

[0167] Sending photos and audio data

[0168] The user sends the captured photo and audio data to the server via the device, which then packages the data and transfers it to the server via the network.

[0169] Input: Photo data, audio data

[0170] Output: Data transferred to the server

[0171] Step 3:

[0172] Data storage and preprocessing

[0173] The server temporarily stores the received photo and audio data, then analyzes the photo metadata (date and time of shooting, location, etc.) and preprocesses the audio data.

[0174] Input: Data transferred to the server

[0175] Output: Stored photo data, audio data, analyzed metadata

[0176] Step 4:

[0177] Automatic photo processing and compositing

[0178] The server applies AI algorithms to the stored photo data, automatically adjusting brightness and contrast, applying filters, and sometimes even combining multiple photos.

[0179] Input: Saved photo data

[0180] Output: Processed and composited photo data

[0181] Step 5:

[0182] Identifying and displaying relevant photo memories

[0183] The server searches a database to identify past photos relevant to that day and displays these photos along with a special message to the user.

[0184] Input: Parsed metadata

[0185] Output: Display memorable photos, special messages

[0186] Step 6:

[0187] Virtual avatar generation

[0188] The server uses the user's photo and voice data to apply DeepFake technology to generate a digital avatar that resembles the user, which is then used to assist with the try-on experience in the virtual store.

[0189] Input: Saved photo data, audio data

[0190] Output: Digital avatar

[0191] Step 7:

[0192] Try-on support in a virtual store

[0193] The user's digital avatar is used to simulate trying on products in a virtual store, allowing the user to see in real time how the selected products will fit and look.

[0194] Input: Digital avatar, digital product data

[0195] Output: Virtual try-on simulation results

[0196] This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home.

[0197] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0198] The present invention provides an advanced photo editing and AI dialogue system that utilizes user photo and voice data. By combining the system with an emotion engine, the system can recognize user emotions and improve the quality of dialogue and photo editing. The system of the present invention is comprised of a server, a terminal, and a user, with each element playing a different role.

[0199] First, a user takes a photo using a device and sends it to a server. The device provides an interface for sending the photo data and its metadata (such as the date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken during a family trip from the device to the server. This operation causes the server to receive the photos and their metadata.

[0200] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0201] In addition, the server identifies past memorable photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0202] Next, the system incorporates an emotion engine to recognize the user's emotions. The server analyzes the user's voice and facial expression data to identify the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy" and influences the behavior of the entire system.

[0203] The emotion engine dynamically changes the photo processing and filter selection depending on the user's emotional state: for example, if the user is sad, it applies a soothing warm-toned filter, while if the user is happy, it adds vibrant colors and special effects to make the photo even more appealing.

[0204] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, a user can provide their own photo and voice to the server, which then uses DeepFake technology to generate an AI model that looks exactly like the user. An emotion engine also influences this conversation, generating responses based on the user's emotions, making the interaction experience more natural and pleasant.

[0205] This system will enable users to automatically enhance their photos, display memorable photos related to special occasions, and use DeepFake technology and an emotion engine to engage in real-time emotional dialogue with an AI model that resembles them, improving the efficiency of photo management and use and the overall user experience.

[0206] The processing flow will be explained below.

[0207] Step 1:

[0208] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0209] Step 2:

[0210] The device sends the selected photos and metadata to the server, where the actual process of uploading data takes place using Wi-Fi or mobile data.

[0211] Step 3:

[0212] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0213] Step 4:

[0214] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0215] Step 5:

[0216] The server stores the AI-processed photos in a final storage area, along with the original metadata, making them easier to search and view in the future.

[0217] Step 6:

[0218] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0219] Step 7:

[0220] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0221] Step 8:

[0222] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0223] Step 9:

[0224] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0225] Step 10:

[0226] The server uses an emotion engine to analyze the user's emotions, which involves analyzing the user's voice and facial expression data in real time to identify their emotional state. For example, if the user's voice is high-pitched, the emotion engine will recognize "tension."

[0227] Step 11:

[0228] Based on the analysis results of the emotion engine, the server dynamically changes the photo processing and filter selection. For example, if the user is analyzed as feeling depressed, a warm-toned filter that gives a sense of comfort is applied.

[0229] Step 12:

[0230] The user initiates a dialogue with the AI ​​through their device, and the server receives input from the user in real time and generates a response using an AI model based on information from the emotion engine. For example, if the user says, "I'm tired today," the AI ​​will suggest, "Thank you for your hard work. How about this filter to help you relax?"

[0231] Through these steps, the system will be able to automatically edit photos, identify and display memorable photos from the past, and realize real-time dialogue using Deep Fake technology and an emotion engine.

[0232] Example 2

[0233] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0234] Conventional photo editing and AI dialogue systems do not take user emotions into account, resulting in a consistent quality of dialogue and photo editing, limiting the user experience. Furthermore, due to a lack of functionality for effectively utilizing past memorable photos, there was a need for a more comprehensive function for automatically displaying memories related to a specific day. Furthermore, there were insufficient means for realizing natural dialogue in real time based on the user's photos and voice.

[0235] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0236] In this invention, the server includes a means for receiving photos, a means for automatically processing and combining the received photos, a means for identifying and displaying photos associated with memories of that day, a means for engaging in real-time dialogue with an AI generated based on the user's photos and voice, and a means for recognizing the user's emotions and dynamically changing the method of processing the photos according to the emotions. This enables high-quality photo processing and natural dialogue according to the user's emotions, and effectively displays memories associated with a specific day.

[0237] "Means for receiving photos" refers to the function of sending photo data taken by a user via a terminal to a server and receiving that data.

[0238] "Means for automatically processing and combining received photographs" refers to a function that uses an artificial intelligence algorithm to adjust the brightness, contrast, and color tone of received photographs, apply various filters as needed, and combine multiple photographs.

[0239] "Means for identifying and displaying photos that have memories related to that day" refers to the function of searching a database for past photos taken on a specific day and notifying and displaying them to the user.

[0240] "Means for engaging in real-time dialogue with AI generated based on the user's photos and voice" refers to a function that allows users to engage in natural dialogue in real time with a digital agent generated using Deep Fake technology based on the photos and voice data provided by the user.

[0241] "Means of recognizing the user's emotions and dynamically changing the way photos are edited depending on the emotions" refers to a function that uses an emotion engine to analyze the user's voice and facial expression data and dynamically change the way photos are edited depending on their emotional state.

[0242] This invention is a system that enables advanced photo processing and dialogue with AI using photo and voice data. This system is composed of a server, a terminal, and a user, and each element functions in cooperation with the others.

[0243] Hardware and software used

[0244] Server: A high-performance server computer responsible for storing photo and audio data, running AI algorithms, and emotion recognition. It uses cloud storage (e.g., Amazon S3, Google® Cloud Storage), AI algorithms (e.g., OpenCV, TENSORFLOW®), emotion engines (e.g., Amazon Rekognition, Microsoft® Azure® Cognitive Services), and Deep Fake technology (e.g., DeepFaceLab, StyleGAN).

[0245] Device: A smartphone, tablet, computer, etc. that provides an interface for taking photos, recording audio, and transmitting data.

[0246] User: Takes photos, inputs voice, sends and receives data, and interacts with AI.

[0247] Data processing and calculation

[0248] 1. Take and send a photo:

[0249] A user takes a photo using a device and sends it to a server via a dedicated application or browser. For example, a user takes a photo of a scene from a family trip with a smartphone and uploads the photo to a server.

[0250] 2. Photo data storage and processing:

[0251] The server stores the received photos in cloud storage. At the same time, it automatically processes the photos using AI algorithms. Specifically, it uses OpenCV and TensorFlow to adjust brightness, contrast, and color tone, apply filters, and combine multiple photos. For example, the server analyzes a landscape photo and adjusts color tone and contrast to create a clearer image.

[0252] 3. Identify and view photos of past memories:

[0253] The server searches the database for past photos related to a specific date and notifies and displays them to the user. This is done using SQL queries. For example, on a user's birthday, photos taken on past birthdays can be displayed with a special message.

[0254] 4. Emotion recognition and dynamic photo manipulation:

[0255] The server analyzes the user's voice and facial expression data and uses an emotion engine to recognize the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy." The server then dynamically changes the photo editing method and filter selection based on the emotion engine's recognition results. For example, if the user is sad, the server applies a warm-toned filter to edit the photo to match the user's emotion.

[0256] 5. Creating and interacting with AI models using Deep Fake technology:

[0257] The user provides the server with their own photo and voice data. Based on this data, the server uses Deep Fake technology to generate a digital agent that looks exactly like the user. The generated digital agent can then engage in natural conversations with the user in real time. For example, if the user says, "I'm very happy today," the agent will respond, "That's great. Why are you happy?"

[0258] Examples and prompts

[0259] Examples:

[0260] When a user sends photos taken on a family trip from their device to the server, the server automatically processes the photos and displays memories of the trip. The photo processing also changes depending on the user's emotions. Finally, a digital agent generated based on the user's photos and voice data interacts with the user in real time, and the interaction changes depending on the user's emotions.

[0261] Example prompt:

[0262] Submit photos from your family vacation and create an AI model that produces clearer images.

[0263] Search for photos from past birthdays and display them to the user with a special message.

[0264] Analyze a user's photo and voice data, generate a digital agent that resembles the user in real time using Deep Fake technology, and generate responses based on the user's emotions.

[0265] The above is the details regarding the "Description of the Preferred Embodiments."

[0266] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0267] Specific explanation of program processing

[0268] Processing Steps

[0269] Step 1:

[0270] A user takes a photo using a device and sends it to a server. The input data is the photo file and its metadata (date and time of the photo, location, etc.), and the output data is the photo file transferred to the server. In concrete terms, a user takes photos of scenery from a family trip with their smartphone and uploads the photos to the server via a dedicated application.

[0271] Step 2:

[0272] The server temporarily stores the received photo data in cloud storage. The input data is the photo file and metadata sent from the device, and the output data is the photo file stored in cloud storage. Specifically, the server stores the photo data and metadata in Amazon S3 or Google Cloud Storage.

[0273] Step 3:

[0274] The server automatically processes stored photos using AI algorithms. The input data is the stored photo file, and the output data is the processed photo file. Specifically, the server uses OpenCV and TensorFlow to adjust the brightness, contrast, and color tone of the photo and apply filters as needed. For example, it adjusts the color tone and contrast of a landscape photo to create a clearer image.

[0275] Step 4:

[0276] The server identifies past memorable photos related to that day from the database and displays them to the user. The input data is the user's photo database and the current date information, and the output data is the identified past memorable photos. Specifically, the server uses an SQL query to search for past photos taken on a specific day and notifies and displays them to the user. For example, on the user's birthday, photos taken on past birthdays are displayed with a special message.

[0277] Step 5:

[0278] The server analyzes the user's voice and facial expression data and recognizes the user's emotions using an emotion engine. The input data is the user's voice file and facial expression image, and the output data is the recognized emotional state. Specifically, the server analyzes the voice and facial expression using an emotion engine (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services) to identify the user's emotional state. For example, if the user is smiling, it is recognized as "joy."

[0279] Step 6:

[0280] The server dynamically changes the photo processing method according to the emotion. The input data is the recognized emotional state and the photo file, and the output data is the photo file processed according to the emotion. Specifically, the server uses the results of the emotion engine to apply a warm filter to the photo if sadness is recognized, and add a vivid filter or special effect if joy is recognized.

[0281] Step 7:

[0282] The server generates an AI model using DeepFake technology based on the user's photo and voice data. The input data is the user's photo and voice files, and the output data is the generated digital agent. Specifically, the server uses DeepFaceLab and StyleGAN to generate a digital agent that looks exactly like the user.

[0283] Step 8:

[0284] The user interacts with the generated digital agent in real time. The input data is the user's real-time voice and text input, and the output data is the digital agent's response. Specifically, when the user says, "I'm very happy today," the agent responds, "That's great. Why are you happy?", and the interaction progresses naturally.

[0285] The above is a concrete explanation of the processing contents of the program of this system.

[0286] (Application example 2)

[0287] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0288] Existing photo editing systems lack the ability to take user emotions into account and provide AI-based interactive dialogue. In particular, it is difficult to provide personalized responses based on photos and voice, limiting the user experience. Furthermore, the lack of a function to automatically identify and display memorable photos from the past on special occasions leaves users without a way to share their memories more effectively.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0290] In this invention, the server includes a means for automatically processing and synthesizing photos, a means for recognizing a user's emotions and improving the quality of photo processing and dialogue in response to the emotions, and a means for generating a digital agent using DeepFake technology based on the user's photo and voice data. This enables photo processing in response to the user's emotions, enabling interactive dialogue that takes emotions into account. Furthermore, by identifying past memorable photos and displaying them on special days, the user's memory sharing experience can be improved.

[0291] The "means for receiving photos" is a function for receiving photo data and its metadata sent from the user's terminal.

[0292] "Means for automatically processing and compositing received photos" refers to a function that uses an AI algorithm to adjust the brightness, contrast, and color tone of received photos, apply filters, and composite multiple photos.

[0293] The "means for identifying and displaying photos with memories related to that day" is a function for identifying past photos related to a date specified by the user and displaying those photos to the user.

[0294] "Means of generating AI based on the user's photo and voice and holding a conversation" refers to a function that uses the user's photo and voice data to generate a digital agent that resembles the user, and that agent then holds a conversation with the user.

[0295] "Means to recognize the user's emotions and improve the quality of photo editing and dialogue according to the emotions" refers to a function that analyzes the user's facial expressions and voice to identify their emotional state, and applies appropriate filters according to that emotional state, improving the quality of dialogue with AI.

[0296] "Means for generating a digital agent resembling a user using DeepFake technology" refers to a function that uses DeepFake technology to create a digital agent resembling a user based on photographs and audio data provided by the user.

[0297] This invention provides an advanced photo processing and AI dialogue system that utilizes user photo and voice data. This system realizes photo reception, emotion recognition, photo processing, specific display of past photos, and interactive dialogue through the application of DeepFake technology.

[0298] Program Description

[0299] Receiving photos

[0300] First, a user takes a photo using their own device and sends it to the server. For this operation, the device provides an interface for sending photo data and its metadata (date and time of shooting, location, etc.) to the server. For example, when a user sends photos taken during a family trip from their device to the server, the server receives the photos and metadata and temporarily stores them.

[0301] emotion recognition

[0302] The server then analyzes the received photos and user-provided audio data, using emotion recognition libraries such as DeepFace and EmotionRecognizer to identify the user's emotional state from their facial expressions and vocal tone. For example, if the user is smiling in the photo, the system will recognize this as "joy."

[0303] Photo editing

[0304] Based on the emotion recognition results, the server automatically processes the photo, including adjusting brightness and contrast, applying filters, and even merging photos. Using AI algorithms, the server applies the optimal processing based on the user's emotional state. For example, if the user is smiling, a vibrant filter is applied, while if not, a warmer color tone is applied.

[0305] Specific display of past photos

[0306] The server also identifies past photo memories associated with the specified date. The feature searches the user's database of photos to find those taken on that date. It notifies the user with a special message, allowing them to share memories more deeply. For example, when a user celebrates their birthday, it presents photos taken on past birthdays.

[0307] Interacting with a digital agent

[0308] Finally, a digital agent is generated using DeepFake technology based on the user's photo and voice data. This allows the user to interact with a digital agent that resembles them in real time. Emotion recognition results are also reflected in the dialogue, and responses are generated that reflect the user's emotions.

[0309] Examples of concrete examples and prompts

[0310] A user takes a photo of a family trip with their smartphone and sends a prompt to the system: "Analyze the emotions of the people in this photo and apply the most appropriate filter." The system recognizes the emotion and returns the photo with a vivid filter applied. Alternatively, by entering a prompt such as "Identify the emotion in this voice and decide which filter to apply to the photo and which story to generate," the system can identify emotions based on the voice data and create appropriate photo editing and story generation.

[0311] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0312] Step 1:

[0313] A user takes a photo on their device and sends the photo and metadata (date, time, location, etc.) from the device to the server. This input includes image data and metadata, which the server receives and temporarily stores. Specifically, the file is uploaded through the device's interface.

[0314] Step 2:

[0315] The server analyzes the received photos and audio data sent by the user. It uses emotion recognition software such as DeepFace or EmotionRecognizer to identify emotions from facial expressions and voice. The input includes image data and audio data, and the output is the emotion recognition results. Specific operations include image processing and audio analysis.

[0316] Step 3:

[0317] The server automatically processes the received photos based on the emotion recognition results. It uses AI algorithms (such as the Python library OpenCV) to adjust brightness, contrast, and color tone, and apply filters. The input includes the original image data and the emotion recognition results, and the output is a processed photo. Specifically, image processing is performed to change the numerical data of the image.

[0318] Step 4:

[0319] The server searches the user's database for photos of past memories related to the specified date and identifies them. The input includes the user's photo database and the specified date, and the output is the identification and extraction of the relevant photos. Specifically, a database search algorithm is executed.

[0320] Step 5:

[0321] The server displays the identified memorable photo to the user. During this process, a special message is attached and notified to the user's device. The input includes the identified memorable photo and the message, and the output is a notification to the user's device. The specific operation is to send the message using the notification system.

[0322] Step 6:

[0323] The server uses DeepFake technology to generate a digital agent based on the user's photo and voice data. The input includes image and voice data, and the generated digital agent is obtained as the output. Specifically, DeepFake modeling and synthesis processing are performed.

[0324] Step 7:

[0325] The user interacts with the generated digital agent in real time. During this process, emotion recognition results are reflected in the dialogue, resulting in more natural responses. The input includes the user's real-time voice data and emotion recognition results, and the digital agent's response is generated as output. Specific operations include natural language processing and real-time speech synthesis.

[0326] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0327] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0328] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0329] [Second embodiment]

[0330] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0331] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0332] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0333] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0334] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0335] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0336] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0337] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0338] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0339] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0340] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0341] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0342] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0343] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0344] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0345] In addition, the server identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. For example, if a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0346] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, if the user provides their own photo and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to enjoy a conversation with the AI.

[0347] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them using Deep Fake technology, improving the efficiency of photo management and use and enhancing the user experience.

[0348] The processing flow will be explained below.

[0349] Step 1:

[0350] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0351] Step 2:

[0352] The device sends the selected photos and metadata to the server, where the actual uploading process takes place using Wi-Fi or mobile data.

[0353] Step 3:

[0354] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0355] Step 4:

[0356] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0357] Step 5:

[0358] The server stores the AI-processed photos in a final storage area, along with the original metadata, making future searches and display easier.

[0359] Step 6:

[0360] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0361] Step 7:

[0362] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0363] Step 8:

[0364] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0365] Step 9:

[0366] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0367] Step 10:

[0368] The user initiates a dialogue with the AI ​​through the device, and the server receives input from the user in real time and generates a response using the AI ​​model. For example, if you ask, "What's the weather like today?", the AI ​​will provide weather information in the format desired by the user. This dialogue can be done via text or voice, improving the user experience.

[0369] Through these steps, the system enables automatic photo editing, identification and display of memorable photos, and real-time interaction using Deep Fake technology.

[0370] Example 1

[0371] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0372] In recent years, users have been taking and storing a large number of digital photos, which requires time and effort to organize and edit them. Furthermore, it is difficult to easily look back on past photos of memories, creating a demand for new ways of communicating using photos. Furthermore, while there is hope for the generation and use of AI models that can converse in real time using the user's own photos and voice, achieving this requires advanced technology.

[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0374] In this invention, the server includes means for taking and sending photos from a terminal, means for receiving and storing the photos and their metadata on the server, means for automatically processing and synthesizing the received photos using an AI algorithm, means for searching and displaying past memorable photos related to that day from a database, and means for generating an AI model using DeepFake technology based on the user's photos and voice and conducting real-time conversations with the user. This allows users to automate the organization and processing of photos, easily look back on memorable photos from special days, and enjoy a new communication experience of interacting in real time with an AI model that resembles them.

[0375] A "terminal" is an electronic device that a user uses to take a photo and transmit the captured photo data and its metadata to a server.

[0376] "Photo" refers to image data that is taken by a user using a terminal and sent to a server.

[0377] "Metadata" is auxiliary information that accompanies photo data, such as the date and time the photo was taken and the location where it was taken.

[0378] The "server" is a system that receives photo data and metadata sent from the device, temporarily stores them, and then processes and synthesizes this data using AI algorithms.

[0379] An "AI algorithm" is an artificial intelligence calculation method used by the server to automatically process and synthesize photo data.

[0380] "Automatic processing" is a process that uses AI algorithms to automatically edit received photo data, such as adjusting brightness, contrast, and applying filters.

[0381] "Synthesis" is the process of combining multiple photographic data to generate a new image.

[0382] "Memorable photos" are specific photos taken by the user in the past that are related to that day.

[0383] "Deep Fake technology" is a technology that generates a realistic AI model that resembles a user based on the user's photo and voice data.

[0384] An "AI model" is a digital agent that resembles a user and is generated using Deep Fake technology based on the user's photo and voice data.

[0385] "Real-time conversation" is a form of communication in which a user interacts with a generated AI model and receives an immediate response.

[0386] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0387] First, a user takes a photo using a device and sends it to a server. A dedicated application is installed on the device, and the user uploads the photo using this application. The device provides a transmission function that includes photo data and its metadata (date and time of photo, location, etc.). A specific example is when a user takes a photo while traveling and sends it to a server via a dedicated application. For example, a user takes a photo of the Eiffel Tower and sends it to a server using the application.

[0388] The server then temporarily stores the received photos and their metadata. The server then automatically processes and combines the stored data using AI algorithms. Specifically, it uses Python scripts and image processing libraries (e.g., OpenCV) to adjust brightness and contrast, apply filters, or combine multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0389] In addition, the server searches the database for past memory photos related to that day and displays them to the user. This function identifies photos taken on the same day in the past and notifies the user. For example, when a user has a birthday, the server finds photos taken on past birthdays and displays them as "past memories" with a special message.

[0390] Finally, an AI model is generated using DeepFake technology based on the user's photo and voice data, allowing the user to enjoy real-time conversations with a digital agent that resembles them. For example, if a user provides their own face and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to have natural conversations with this AI model.

[0391] An example prompt might be:

[0392] "I want to edit a photo of the Eiffel Tower I took during my trip to make it brighter and with more contrast."

[0393] "Find photos from past birthdays and display them with special messages."

[0394] "I would like to use the photos and voice data I provide to generate an AI model that resembles me and then have a conversation with that AI model."

[0395] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them in real time using Deep Fake technology, which is expected to improve the efficiency of photo organization and editing and enhance the user experience.

[0396] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0397] Step 1:

[0398] The user takes a photo on their device, opens a dedicated application, selects the photo, and presses the send button.

[0399] Input: The photo taken and its metadata (date, timestamp, location, etc.).

[0400] Processing: The device uploads the photo data and metadata to the server via a dedicated application.

[0401] Output: Notification of completion of transmission to the server.

[0402] Specific operation: The user takes a photo of a landscape using the camera app on their smartphone, then opens the dedicated app, taps the "Upload Photo" button, and sends the selected photo using the "Send" button.

[0403] Step 2:

[0404] The server receives the photo data and metadata sent from the device and temporarily stores them.

[0405] Input: Photos and metadata sent from the device.

[0406] Processing: The server receives the HTTP request, analyzes the photo data and metadata, and stores the data in a storage service.

[0407] Output: Notification that the photo data and metadata have been saved to storage.

[0408] Specific operation: The server saves the received photos to a storage service (e.g. Amazon S3). After saving is complete, add a record of the save to the database.

[0409] Step 3:

[0410] The server automatically processes and synthesizes the stored photo data using AI algorithms.

[0411] Input: Saved photo data.

[0412] Processing: The server runs a Python script that uses image processing libraries (e.g. OpenCV) to adjust brightness and contrast, apply filters, and possibly blend multiple photos together.

[0413] Output: Processed photo data.

[0414] What it does: The server processes the stored photos sequentially, applying a sepia filter to a photo of the Eiffel Tower, for example, and adjusting the brightness and contrast of another photo.

[0415] Step 4:

[0416] The server searches the database for past memorable photos related to that day, identifies them, and notifies the user.

[0417] Input: Date and time information for that day.

[0418] Processing: The server executes an SQL query to search the database for previous photos taken on the same day. If any are found, the user is notified.

[0419] Output: Identified past memory photos and notification messages.

[0420] What it does: On the user's birthday, the server finds photos taken on past birthdays and notifies the user via the app with a special message.

[0421] Step 5:

[0422] The user sends their own photo and voice data to the server.

[0423] Input: User photo and voice data.

[0424] Processing: The device uploads the photo and audio data to the server via a dedicated application.

[0425] Output: Notification of completion of transmission to the server.

[0426] Specific operation: The user takes a photo of their face and records their voice using a dedicated app, and then sends the data to the server.

[0427] Step 6:

[0428] Based on the received data, the server uses Deep Fake technology to generate an AI model that resembles the user.

[0429] Input: User photo and voice data.

[0430] Processing: The server uses a DeepFake generation tool (e.g., DeepFaceLab) to generate an AI model, which is synthesized based on the user's photo and voice.

[0431] Output: The generated AI model.

[0432] How it works: The server uses deep learning to generate a realistic AI model based on the user's photo and voice data. This AI model is stored on the server and used to interact with the user.

[0433] Step 7:

[0434] Users interact with the generated AI model in real time.

[0435] Input: The generated AI model.

[0436] Processing: The user initiates a conversation with the AI ​​model using a dedicated app. The server runs the necessary infrastructure (e.g., a chat engine) and responds immediately to the user's questions.

[0437] Output: Interaction logs and responses with the AI ​​model.

[0438] How it works: Users communicate with the AI ​​model through text and voice chat. For example, they can ask, "What's the weather like today?" and the AI ​​model will respond in real time.

[0439] Through these steps, users can automatically enhance their photos, easily relive memorable photos from special occasions, and even interact with an AI model that resembles them in real time.

[0440] (Application example 1)

[0441] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0442] Conventional photo editing systems and conversational AI systems have limited advanced interaction and utilization using user photos and voice data, particularly in the virtual store try-on experience. The present invention aims to solve these problems by providing a system that uses user photos and voice data to achieve more advanced photo editing and natural dialogue with a digital agent, and further improves the virtual store try-on experience.

[0443] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0444] In this invention, the server includes means for receiving photos, means for automatically processing and combining the received photos, means for identifying and displaying photos that are memorable for that day, means for generating an AI based on the user's photos and voice to hold a conversation, and means for generating a virtual avatar based on the user's photos and voice to assist in trying on products in a virtual store. This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home, allowing them to check the look and fit of products.

[0445] "Means for receiving photos" means any device or software for sending and receiving photos and their metadata from a user's terminal to the server.

[0446] "Means for automatically processing and synthesizing received photographs" refers to devices or software that use AI algorithms on received photographic data to automatically adjust brightness, contrast, color tone, and apply filters.

[0447] "Means for identifying and displaying memorable photographs relating to that date" refers to a device or software for searching a database for past photographs relating to a particular date and displaying them to a user.

[0448] "Means for generating AI based on a user's photo and voice to converse" refers to a device or software that uses photo and voice data obtained from a user to generate a digital agent that resembles the user using Deep Fake technology, allowing the user to converse naturally with that agent.

[0449] "Means for generating a virtual avatar and assisting in trying on products in a virtual store" refers to a device or software that generates a digital avatar based on a user's photograph and voice, and enables the user to simulate trying on products in a virtual store using that avatar.

[0450] This invention provides an advanced photo processing and AI dialogue system that utilizes photos and voice data taken by users. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0451] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0452] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms. This includes adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by the user and adjusts color and contrast to create a clearer image. The server also identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on that day and displays them with a special message.

[0453] Another distinctive feature of this invention is the ability to use DeepFake technology to generate a digital avatar based on the user's photo and voice data, helping them try on products in a virtual store. Users take photos of themselves with their smartphones and provide voice data. This generates a digital avatar, allowing them to try on products in the virtual store. The generated digital avatar allows the user to try on selected products in real time and check their appearance and fit.

[0454] Hardware and software used:

[0455] Smartphone or PC (used for capturing photos and audio)

[0456] Camera (takes photos)

[0457] Microphone (acquires audio data)

[0458] Server (temporary storage of photo and audio data, processing, generation of virtual avatars)

[0459] OpenCV (used for capturing and displaying photos)

[0460] Deep Fake technology (used to generate digital avatars)

[0461] Python script (used to coordinate the entire process)

[0462] Examples:

[0463] Let's say a user wants to try on new clothes. They launch the app and capture their photo and audio data. Based on this, the app generates a digital avatar and shows them how the selected clothes will look in real time. Through virtual try-on, users can see how the products will look at home without having to go to a physical store.

[0464] Example prompt sentence:

[0465] Prompt: "Generate a digital agent that resembles the user based on the following photo and audio data. The photo and audio data are attached. This agent will act as an avatar for the user to try on products in a virtual store."

[0466] This approach allows users to enjoy a virtual store experience and avoids the hassle of visiting a physical store.

[0467] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0468] Step 1:

[0469] Photo and audio data capture

[0470] Users can take their own photos and record audio data using a smartphone or PC, using the camera to capture the photos and the microphone to capture the audio data.

[0471] Input: Smartphone, PC, camera, microphone

[0472] Output: Photo data, audio data

[0473] Step 2:

[0474] Sending photos and audio data

[0475] The user sends the captured photo and audio data to the server via the device, which then packages the data and transfers it to the server via the network.

[0476] Input: Photo data, audio data

[0477] Output: Data transferred to the server

[0478] Step 3:

[0479] Data storage and preprocessing

[0480] The server temporarily stores the received photo and audio data, then analyzes the photo metadata (date and time of shooting, location, etc.) and preprocesses the audio data.

[0481] Input: Data transferred to the server

[0482] Output: Stored photo data, audio data, analyzed metadata

[0483] Step 4:

[0484] Automatic photo processing and compositing

[0485] The server applies AI algorithms to the stored photo data, automatically adjusting brightness and contrast, applying filters, and sometimes even combining multiple photos.

[0486] Input: Saved photo data

[0487] Output: Processed and composited photo data

[0488] Step 5:

[0489] Identifying and displaying relevant photo memories

[0490] The server searches a database to identify past photos relevant to that day and displays these photos along with a special message to the user.

[0491] Input: Parsed metadata

[0492] Output: Display memorable photos, special messages

[0493] Step 6:

[0494] Virtual avatar generation

[0495] The server uses the user's photo and voice data to apply DeepFake technology to generate a digital avatar that resembles the user, which is then used to assist with the try-on experience in the virtual store.

[0496] Input: Saved photo data, audio data

[0497] Output: Digital avatar

[0498] Step 7:

[0499] Try-on support in a virtual store

[0500] The user's digital avatar is used to simulate trying on products in a virtual store, allowing the user to see in real time how the selected products will fit and look.

[0501] Input: Digital avatar, digital product data

[0502] Output: Virtual try-on simulation results

[0503] This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home.

[0504] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0505] The present invention provides an advanced photo editing and AI dialogue system that utilizes user photo and voice data. By combining the system with an emotion engine, the system can recognize user emotions and improve the quality of dialogue and photo editing. The system of the present invention is comprised of a server, a terminal, and a user, with each element playing a different role.

[0506] First, a user takes a photo using a device and sends it to a server. The device provides an interface for sending the photo data and its metadata (such as the date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken during a family trip from the device to the server. This operation causes the server to receive the photos and their metadata.

[0507] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0508] In addition, the server identifies past memorable photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0509] Next, the system incorporates an emotion engine to recognize the user's emotions. The server analyzes the user's voice and facial expression data to identify the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy" and influences the behavior of the entire system.

[0510] The emotion engine dynamically changes the photo processing and filter selection depending on the user's emotional state: for example, if the user is sad, it applies a soothing warm-toned filter, while if the user is happy, it adds vibrant colors and special effects to make the photo even more appealing.

[0511] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, a user can provide their own photo and voice to the server, which then uses DeepFake technology to generate an AI model that looks exactly like the user. An emotion engine also influences this conversation, generating responses based on the user's emotions, making the interaction experience more natural and pleasant.

[0512] This system will enable users to automatically enhance their photos, display memorable photos related to special occasions, and use DeepFake technology and an emotion engine to engage in real-time emotional dialogue with an AI model that resembles them, improving the efficiency of photo management and use and the overall user experience.

[0513] The processing flow will be explained below.

[0514] Step 1:

[0515] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0516] Step 2:

[0517] The device sends the selected photos and metadata to the server, where the actual process of uploading data takes place using Wi-Fi or mobile data.

[0518] Step 3:

[0519] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0520] Step 4:

[0521] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0522] Step 5:

[0523] The server stores the AI-processed photos in a final storage area, along with the original metadata, making them easier to search and view in the future.

[0524] Step 6:

[0525] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0526] Step 7:

[0527] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0528] Step 8:

[0529] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0530] Step 9:

[0531] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0532] Step 10:

[0533] The server uses an emotion engine to analyze the user's emotions, which involves analyzing the user's voice and facial expression data in real time to identify their emotional state. For example, if the user's voice is high-pitched, the emotion engine will recognize "tension."

[0534] Step 11:

[0535] Based on the analysis results of the emotion engine, the server dynamically changes the photo processing and filter selection. For example, if the user is analyzed as feeling depressed, a warm-toned filter that gives a sense of comfort is applied.

[0536] Step 12:

[0537] The user initiates a dialogue with the AI ​​through their device, and the server receives input from the user in real time and generates a response using an AI model based on information from the emotion engine. For example, if the user says, "I'm tired today," the AI ​​will suggest, "Thank you for your hard work. How about this filter to help you relax?"

[0538] Through these steps, the system will be able to automatically edit photos, identify and display memorable photos from the past, and realize real-time dialogue using Deep Fake technology and an emotion engine.

[0539] Example 2

[0540] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0541] Conventional photo editing and AI dialogue systems do not take user emotions into account, resulting in a consistent quality of dialogue and photo editing, limiting the user experience. Furthermore, due to a lack of functionality for effectively utilizing past memorable photos, there was a need for a more comprehensive function for automatically displaying memories related to a specific day. Furthermore, there were insufficient means for realizing natural dialogue in real time based on the user's photos and voice.

[0542] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0543] In this invention, the server includes a means for receiving photos, a means for automatically processing and combining the received photos, a means for identifying and displaying photos associated with memories of that day, a means for engaging in real-time dialogue with an AI generated based on the user's photos and voice, and a means for recognizing the user's emotions and dynamically changing the method of processing the photos according to the emotions. This enables high-quality photo processing and natural dialogue according to the user's emotions, and effectively displays memories associated with a specific day.

[0544] "Means for receiving photos" refers to the function of sending photo data taken by a user via a terminal to a server and receiving that data.

[0545] "Means for automatically processing and combining received photographs" refers to a function that uses an artificial intelligence algorithm to adjust the brightness, contrast, and color tone of received photographs, apply various filters as needed, and combine multiple photographs.

[0546] "Means for identifying and displaying photos that have memories related to that day" refers to the function of searching a database for past photos taken on a specific day and notifying and displaying them to the user.

[0547] "Means for engaging in real-time dialogue with AI generated based on the user's photos and voice" refers to a function that allows users to engage in natural dialogue in real time with a digital agent generated using Deep Fake technology based on the photos and voice data provided by the user.

[0548] "Means of recognizing the user's emotions and dynamically changing the way photos are edited depending on the emotions" refers to a function that uses an emotion engine to analyze the user's voice and facial expression data and dynamically change the way photos are edited depending on their emotional state.

[0549] This invention is a system that enables advanced photo processing and dialogue with AI using photo and voice data. This system is composed of a server, a terminal, and a user, and each element functions in cooperation with the others.

[0550] Hardware and software used

[0551] Server: A high-performance server computer responsible for storing photo and audio data, running AI algorithms, and emotion recognition. It uses cloud storage (e.g., Amazon S3, Google Cloud Storage), AI algorithms (e.g., OpenCV, TensorFlow), emotion engines (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services), and Deep Fake technology (e.g., DeepFaceLab, StyleGAN).

[0552] Device: A smartphone, tablet, computer, etc. that provides an interface for taking photos, recording audio, and transmitting data.

[0553] User: Takes photos, inputs voice, sends and receives data, and interacts with AI.

[0554] Data processing and calculation

[0555] 1. Take and send a photo:

[0556] A user takes a photo using a device and sends it to a server via a dedicated application or browser. For example, a user takes a photo of a scene from a family trip with a smartphone and uploads the photo to a server.

[0557] 2. Photo data storage and processing:

[0558] The server stores the received photos in cloud storage. At the same time, it automatically processes the photos using AI algorithms. Specifically, it uses OpenCV and TensorFlow to adjust brightness, contrast, and color tone, apply filters, and combine multiple photos. For example, the server analyzes a landscape photo and adjusts color tone and contrast to create a clearer image.

[0559] 3. Identify and view photos of past memories:

[0560] The server searches the database for past photos related to a specific date and notifies and displays them to the user. This is done using SQL queries. For example, on a user's birthday, photos taken on past birthdays can be displayed with a special message.

[0561] 4. Emotion recognition and dynamic photo manipulation:

[0562] The server analyzes the user's voice and facial expression data and uses an emotion engine to recognize the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy." The server then dynamically changes the photo editing method and filter selection based on the emotion engine's recognition results. For example, if the user is sad, the server applies a warm-toned filter to edit the photo to match the user's emotion.

[0563] 5. Creating and interacting with AI models using Deep Fake technology:

[0564] The user provides the server with their own photo and voice data. Based on this data, the server uses Deep Fake technology to generate a digital agent that looks exactly like the user. The generated digital agent can then engage in natural conversations with the user in real time. For example, if the user says, "I'm very happy today," the agent will respond, "That's great. Why are you happy?"

[0565] Examples and prompts

[0566] Examples:

[0567] When a user sends photos taken on a family trip from their device to the server, the server automatically processes the photos and displays memories of the trip. The photo processing also changes depending on the user's emotions. Finally, a digital agent generated based on the user's photos and voice data interacts with the user in real time, and the interaction changes depending on the user's emotions.

[0568] Example prompt:

[0569] Submit photos from your family vacation and create an AI model that produces clearer images.

[0570] Search for photos from past birthdays and display them to the user with a special message.

[0571] Analyze a user's photo and voice data, generate a digital agent that resembles the user in real time using Deep Fake technology, and generate responses based on the user's emotions.

[0572] The above is the details regarding the "Description of the Preferred Embodiments."

[0573] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0574] Specific explanation of program processing

[0575] Processing Steps

[0576] Step 1:

[0577] A user takes a photo using a device and sends it to a server. The input data is the photo file and its metadata (date and time of the photo, location, etc.), and the output data is the photo file transferred to the server. In concrete terms, a user takes photos of scenery from a family trip with their smartphone and uploads the photos to the server via a dedicated application.

[0578] Step 2:

[0579] The server temporarily stores the received photo data in cloud storage. The input data is the photo file and metadata sent from the device, and the output data is the photo file stored in cloud storage. Specifically, the server stores the photo data and metadata in Amazon S3 or Google Cloud Storage.

[0580] Step 3:

[0581] The server automatically processes stored photos using AI algorithms. The input data is the stored photo file, and the output data is the processed photo file. Specifically, the server uses OpenCV and TensorFlow to adjust the brightness, contrast, and color tone of the photo and apply filters as needed. For example, it adjusts the color tone and contrast of a landscape photo to create a clearer image.

[0582] Step 4:

[0583] The server identifies past memorable photos related to that day from the database and displays them to the user. The input data is the user's photo database and the current date information, and the output data is the identified past memorable photos. Specifically, the server uses an SQL query to search for past photos taken on a specific day and notifies and displays them to the user. For example, on the user's birthday, photos taken on past birthdays are displayed with a special message.

[0584] Step 5:

[0585] The server analyzes the user's voice and facial expression data and recognizes the user's emotions using an emotion engine. The input data is the user's voice file and facial expression image, and the output data is the recognized emotional state. Specifically, the server analyzes the voice and facial expression using an emotion engine (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services) to identify the user's emotional state. For example, if the user is smiling, it is recognized as "joy."

[0586] Step 6:

[0587] The server dynamically changes the photo processing method according to the emotion. The input data is the recognized emotional state and the photo file, and the output data is the photo file processed according to the emotion. Specifically, the server uses the results of the emotion engine to apply a warm filter to the photo if sadness is recognized, and add a vivid filter or special effect if joy is recognized.

[0588] Step 7:

[0589] The server generates an AI model using DeepFake technology based on the user's photo and voice data. The input data is the user's photo and voice files, and the output data is the generated digital agent. Specifically, the server uses DeepFaceLab and StyleGAN to generate a digital agent that looks exactly like the user.

[0590] Step 8:

[0591] The user interacts with the generated digital agent in real time. The input data is the user's real-time voice and text input, and the output data is the digital agent's response. Specifically, when the user says, "I'm very happy today," the agent responds, "That's great. Why are you happy?", and the interaction progresses naturally.

[0592] The above is a concrete explanation of the processing contents of the program of this system.

[0593] (Application example 2)

[0594] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0595] Existing photo editing systems lack the ability to take user emotions into account and provide AI-based interactive dialogue. In particular, it is difficult to provide personalized responses based on photos and voice, limiting the user experience. Furthermore, the lack of a function to automatically identify and display memorable photos from the past on special occasions leaves users without a way to share their memories more effectively.

[0596] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0597] In this invention, the server includes a means for automatically processing and synthesizing photos, a means for recognizing a user's emotions and improving the quality of photo processing and dialogue in response to the emotions, and a means for generating a digital agent using DeepFake technology based on the user's photo and voice data. This enables photo processing in response to the user's emotions, enabling interactive dialogue that takes emotions into account. Furthermore, by identifying past memorable photos and displaying them on special days, the user's memory sharing experience can be improved.

[0598] The "means for receiving photos" is a function for receiving photo data and its metadata sent from the user's terminal.

[0599] "Means for automatically processing and compositing received photos" refers to a function that uses an AI algorithm to adjust the brightness, contrast, and color tone of received photos, apply filters, and composite multiple photos.

[0600] The "means for identifying and displaying photos with memories related to that day" is a function for identifying past photos related to a date specified by the user and displaying those photos to the user.

[0601] "Means of generating AI based on the user's photo and voice and holding a conversation" refers to a function that uses the user's photo and voice data to generate a digital agent that resembles the user, and that agent then holds a conversation with the user.

[0602] "Means to recognize the user's emotions and improve the quality of photo editing and dialogue according to the emotions" refers to a function that analyzes the user's facial expressions and voice to identify their emotional state, and applies appropriate filters according to that emotional state, improving the quality of dialogue with AI.

[0603] "Means for generating a digital agent resembling a user using DeepFake technology" refers to a function that uses DeepFake technology to create a digital agent resembling a user based on photographs and audio data provided by the user.

[0604] This invention provides an advanced photo processing and AI dialogue system that utilizes user photo and voice data. This system realizes photo reception, emotion recognition, photo processing, specific display of past photos, and interactive dialogue through the application of DeepFake technology.

[0605] Program Description

[0606] Receiving photos

[0607] First, a user takes a photo using their own device and sends it to the server. For this operation, the device provides an interface for sending photo data and its metadata (date and time of shooting, location, etc.) to the server. For example, when a user sends photos taken during a family trip from their device to the server, the server receives the photos and metadata and temporarily stores them.

[0608] emotion recognition

[0609] The server then analyzes the received photos and user-provided audio data, using emotion recognition libraries such as DeepFace and EmotionRecognizer to identify the user's emotional state from their facial expressions and vocal tone. For example, if the user is smiling in the photo, the system will recognize this as "joy."

[0610] Photo editing

[0611] Based on the emotion recognition results, the server automatically processes the photo, including adjusting brightness and contrast, applying filters, and even merging photos. Using AI algorithms, the server applies the optimal processing based on the user's emotional state. For example, if the user is smiling, a vibrant filter is applied, while if not, a warmer color tone is applied.

[0612] Specific display of past photos

[0613] The server also identifies past photo memories associated with the specified date. The feature searches the user's database of photos to find those taken on that date. It notifies the user with a special message, allowing them to share memories more deeply. For example, when a user celebrates their birthday, it presents photos taken on past birthdays.

[0614] Interacting with a digital agent

[0615] Finally, a digital agent is generated using DeepFake technology based on the user's photo and voice data. This allows the user to interact with a digital agent that resembles them in real time. Emotion recognition results are also reflected in the dialogue, and responses are generated that reflect the user's emotions.

[0616] Examples of concrete examples and prompts

[0617] A user takes a photo of a family trip with their smartphone and sends a prompt to the system: "Analyze the emotions of the people in this photo and apply the most appropriate filter." The system recognizes the emotion and returns the photo with a vivid filter applied. Alternatively, by entering a prompt such as "Identify the emotion in this voice and decide which filter to apply to the photo and which story to generate," the system can identify emotions based on the voice data and create appropriate photo editing and story generation.

[0618] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0619] Step 1:

[0620] A user takes a photo on their device and sends the photo and metadata (date, time, location, etc.) from the device to the server. This input includes image data and metadata, which the server receives and temporarily stores. Specifically, the file is uploaded through the device's interface.

[0621] Step 2:

[0622] The server analyzes the received photos and audio data sent by the user. It uses emotion recognition software such as DeepFace or EmotionRecognizer to identify emotions from facial expressions and voice. The input includes image data and audio data, and the output is the emotion recognition results. Specific operations include image processing and audio analysis.

[0623] Step 3:

[0624] The server automatically processes the received photos based on the emotion recognition results. It uses AI algorithms (such as the Python library OpenCV) to adjust brightness, contrast, and color tone, and apply filters. The input includes the original image data and the emotion recognition results, and the output is a processed photo. Specifically, image processing is performed to change the numerical data of the image.

[0625] Step 4:

[0626] The server searches the user's database for photos of past memories related to the specified date and identifies them. The input includes the user's photo database and the specified date, and the output is the identification and extraction of the relevant photos. Specifically, a database search algorithm is executed.

[0627] Step 5:

[0628] The server displays the identified memorable photo to the user. During this process, a special message is attached and notified to the user's device. The input includes the identified memorable photo and the message, and the output is a notification to the user's device. The specific operation is to send the message using the notification system.

[0629] Step 6:

[0630] The server uses DeepFake technology to generate a digital agent based on the user's photo and voice data. The input includes image and voice data, and the generated digital agent is obtained as the output. Specifically, DeepFake modeling and synthesis processing are performed.

[0631] Step 7:

[0632] The user interacts with the generated digital agent in real time. During this process, emotion recognition results are reflected in the dialogue, resulting in more natural responses. The input includes the user's real-time voice data and emotion recognition results, and the digital agent's response is generated as output. Specific operations include natural language processing and real-time speech synthesis.

[0633] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0634] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0635] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0636] [Third embodiment]

[0637] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0638] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0639] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0640] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0641] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0642] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0643] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0644] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0645] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0646] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0647] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0648] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0649] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0650] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0651] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0652] In addition, the server identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. For example, if a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0653] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, if the user provides their own photo and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to enjoy a conversation with the AI.

[0654] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them using Deep Fake technology, improving the efficiency of photo management and use and enhancing the user experience.

[0655] The processing flow will be explained below.

[0656] Step 1:

[0657] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0658] Step 2:

[0659] The device sends the selected photos and metadata to the server, where the actual uploading process takes place using Wi-Fi or mobile data.

[0660] Step 3:

[0661] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0662] Step 4:

[0663] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0664] Step 5:

[0665] The server stores the AI-processed photos in a final storage area, along with the original metadata, making future searches and display easier.

[0666] Step 6:

[0667] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0668] Step 7:

[0669] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0670] Step 8:

[0671] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0672] Step 9:

[0673] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0674] Step 10:

[0675] The user initiates a dialogue with the AI ​​through the device, and the server receives input from the user in real time and generates a response using the AI ​​model. For example, if you ask, "What's the weather like today?", the AI ​​will provide weather information in the format desired by the user. This dialogue can be done via text or voice, improving the user experience.

[0676] Through these steps, the system enables automatic photo editing, identification and display of memorable photos, and real-time interaction using Deep Fake technology.

[0677] Example 1

[0678] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0679] In recent years, users have been taking and storing a large number of digital photos, which requires time and effort to organize and edit them. Furthermore, it is difficult to easily look back on past photos of memories, creating a demand for new ways of communicating using photos. Furthermore, while there is hope for the generation and use of AI models that can converse in real time using the user's own photos and voice, achieving this requires advanced technology.

[0680] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0681] In this invention, the server includes means for taking and sending photos from a terminal, means for receiving and storing the photos and their metadata on the server, means for automatically processing and synthesizing the received photos using an AI algorithm, means for searching and displaying past memorable photos related to that day from a database, and means for generating an AI model using DeepFake technology based on the user's photos and voice and conducting real-time conversations with the user. This allows users to automate the organization and processing of photos, easily look back on memorable photos from special days, and enjoy a new communication experience of interacting in real time with an AI model that resembles them.

[0682] A "terminal" is an electronic device that a user uses to take a photo and transmit the captured photo data and its metadata to a server.

[0683] "Photo" refers to image data that is taken by a user using a terminal and sent to a server.

[0684] "Metadata" is auxiliary information that accompanies photo data, such as the date and time the photo was taken and the location where it was taken.

[0685] The "server" is a system that receives photo data and metadata sent from the device, temporarily stores them, and then processes and synthesizes this data using AI algorithms.

[0686] An "AI algorithm" is an artificial intelligence calculation method used by the server to automatically process and synthesize photo data.

[0687] "Automatic processing" is a process that uses AI algorithms to automatically edit received photo data, such as adjusting brightness, contrast, and applying filters.

[0688] "Synthesis" is the process of combining multiple photographic data to generate a new image.

[0689] "Memorable photos" are specific photos taken by the user in the past that are related to that day.

[0690] "Deep Fake technology" is a technology that generates a realistic AI model that resembles a user based on the user's photo and voice data.

[0691] An "AI model" is a digital agent that resembles a user and is generated using Deep Fake technology based on the user's photo and voice data.

[0692] "Real-time conversation" is a form of communication in which a user interacts with a generated AI model and receives an immediate response.

[0693] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0694] First, a user takes a photo using a device and sends it to a server. A dedicated application is installed on the device, and the user uploads the photo using this application. The device provides a transmission function that includes photo data and its metadata (date and time of photo, location, etc.). A specific example is when a user takes a photo while traveling and sends it to a server via a dedicated application. For example, a user takes a photo of the Eiffel Tower and sends it to a server using the application.

[0695] The server then temporarily stores the received photos and their metadata. The server then automatically processes and combines the stored data using AI algorithms. Specifically, it uses Python scripts and image processing libraries (e.g., OpenCV) to adjust brightness and contrast, apply filters, or combine multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0696] In addition, the server searches the database for past memory photos related to that day and displays them to the user. This function identifies photos taken on the same day in the past and notifies the user. For example, when a user has a birthday, the server finds photos taken on past birthdays and displays them as "past memories" with a special message.

[0697] Finally, an AI model is generated using DeepFake technology based on the user's photo and voice data, allowing the user to enjoy real-time conversations with a digital agent that resembles them. For example, if a user provides their own face and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to have natural conversations with this AI model.

[0698] An example prompt might be:

[0699] "I want to edit a photo of the Eiffel Tower I took during my trip to make it brighter and with more contrast."

[0700] "Find photos from past birthdays and display them with special messages."

[0701] "I would like to use the photos and voice data I provide to generate an AI model that resembles me and then have a conversation with that AI model."

[0702] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them in real time using Deep Fake technology, which is expected to improve the efficiency of photo organization and editing and enhance the user experience.

[0703] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0704] Step 1:

[0705] The user takes a photo on their device, opens a dedicated application, selects the photo, and presses the send button.

[0706] Input: The photo taken and its metadata (date, timestamp, location, etc.).

[0707] Processing: The device uploads the photo data and metadata to the server via a dedicated application.

[0708] Output: Notification of completion of transmission to the server.

[0709] Specific operation: The user takes a photo of a landscape using the camera app on their smartphone, then opens the dedicated app, taps the "Upload Photo" button, and sends the selected photo using the "Send" button.

[0710] Step 2:

[0711] The server receives the photo data and metadata sent from the device and temporarily stores them.

[0712] Input: Photos and metadata sent from the device.

[0713] Processing: The server receives the HTTP request, analyzes the photo data and metadata, and stores the data in a storage service.

[0714] Output: Notification that the photo data and metadata have been saved to storage.

[0715] Specific operation: The server saves the received photos to a storage service (e.g. Amazon S3). After saving is complete, add a record of the save to the database.

[0716] Step 3:

[0717] The server automatically processes and synthesizes the stored photo data using AI algorithms.

[0718] Input: Saved photo data.

[0719] Processing: The server runs a Python script that uses image processing libraries (e.g. OpenCV) to adjust brightness and contrast, apply filters, and possibly blend multiple photos together.

[0720] Output: Processed photo data.

[0721] What it does: The server processes the stored photos sequentially, applying a sepia filter to a photo of the Eiffel Tower, for example, and adjusting the brightness and contrast of another photo.

[0722] Step 4:

[0723] The server searches the database for past memorable photos related to that day, identifies them, and notifies the user.

[0724] Input: Date and time information for that day.

[0725] Processing: The server executes an SQL query to search the database for previous photos taken on the same day. If any are found, the user is notified.

[0726] Output: Identified past memory photos and notification messages.

[0727] What it does: On the user's birthday, the server finds photos taken on past birthdays and notifies the user via the app with a special message.

[0728] Step 5:

[0729] The user sends their own photo and voice data to the server.

[0730] Input: User photo and voice data.

[0731] Processing: The device uploads the photo and audio data to the server via a dedicated application.

[0732] Output: Notification of completion of transmission to the server.

[0733] Specific operation: The user takes a photo of their face and records their voice using a dedicated app, and then sends the data to the server.

[0734] Step 6:

[0735] Based on the received data, the server uses Deep Fake technology to generate an AI model that resembles the user.

[0736] Input: User photo and voice data.

[0737] Processing: The server uses a DeepFake generation tool (e.g., DeepFaceLab) to generate an AI model, which is synthesized based on the user's photo and voice.

[0738] Output: The generated AI model.

[0739] How it works: The server uses deep learning to generate a realistic AI model based on the user's photo and voice data. This AI model is stored on the server and used to interact with the user.

[0740] Step 7:

[0741] Users interact with the generated AI model in real time.

[0742] Input: The generated AI model.

[0743] Processing: The user initiates a conversation with the AI ​​model using a dedicated app. The server runs the necessary infrastructure (e.g., a chat engine) and responds immediately to the user's questions.

[0744] Output: Interaction logs and responses with the AI ​​model.

[0745] How it works: Users communicate with the AI ​​model through text and voice chat. For example, they can ask, "What's the weather like today?" and the AI ​​model will respond in real time.

[0746] Through these steps, users can automatically enhance their photos, easily relive memorable photos from special occasions, and even interact with an AI model that resembles them in real time.

[0747] (Application example 1)

[0748] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0749] Conventional photo editing systems and conversational AI systems have limited advanced interaction and utilization using user photos and voice data, particularly in the virtual store try-on experience. The present invention aims to solve these problems by providing a system that uses user photos and voice data to achieve more advanced photo editing and natural dialogue with a digital agent, and further improves the virtual store try-on experience.

[0750] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0751] In this invention, the server includes means for receiving photos, means for automatically processing and combining the received photos, means for identifying and displaying photos that are memorable for that day, means for generating an AI based on the user's photos and voice to hold a conversation, and means for generating a virtual avatar based on the user's photos and voice to assist in trying on products in a virtual store. This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home, allowing them to check the look and fit of products.

[0752] "Means for receiving photos" means any device or software for sending and receiving photos and their metadata from a user's terminal to the server.

[0753] "Means for automatically processing and synthesizing received photographs" refers to devices or software that use AI algorithms on received photographic data to automatically adjust brightness, contrast, color tone, and apply filters.

[0754] "Means for identifying and displaying memorable photographs relating to that date" refers to a device or software for searching a database for past photographs relating to a particular date and displaying them to a user.

[0755] "Means for generating AI based on a user's photo and voice to converse" refers to a device or software that uses photo and voice data obtained from a user to generate a digital agent that resembles the user using Deep Fake technology, allowing the user to converse naturally with that agent.

[0756] "Means for generating a virtual avatar and assisting in trying on products in a virtual store" refers to a device or software that generates a digital avatar based on a user's photograph and voice, and enables the user to simulate trying on products in a virtual store using that avatar.

[0757] This invention provides an advanced photo processing and AI dialogue system that utilizes photos and voice data taken by users. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0758] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0759] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms. This includes adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by the user and adjusts color and contrast to create a clearer image. The server also identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on that day and displays them with a special message.

[0760] Another distinctive feature of this invention is the ability to use DeepFake technology to generate a digital avatar based on the user's photo and voice data, helping them try on products in a virtual store. Users take photos of themselves with their smartphones and provide voice data. This generates a digital avatar, allowing them to try on products in the virtual store. The generated digital avatar allows the user to try on selected products in real time and check their appearance and fit.

[0761] Hardware and software used:

[0762] Smartphone or PC (used for capturing photos and audio)

[0763] Camera (takes photos)

[0764] Microphone (acquires audio data)

[0765] Server (temporary storage of photo and audio data, processing, generation of virtual avatars)

[0766] OpenCV (used for capturing and displaying photos)

[0767] Deep Fake technology (used to generate digital avatars)

[0768] Python script (used to coordinate the entire process)

[0769] Examples:

[0770] Let's say a user wants to try on new clothes. They launch the app and capture their photo and audio data. Based on this, the app generates a digital avatar and shows them how the selected clothes will look in real time. Through virtual try-on, users can see how the products will look at home without having to go to a physical store.

[0771] Example prompt sentence:

[0772] Prompt: "Generate a digital agent that resembles the user based on the following photo and audio data. The photo and audio data are attached. This agent will act as an avatar for the user to try on products in a virtual store."

[0773] This approach allows users to enjoy a virtual store experience and avoids the hassle of visiting a physical store.

[0774] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0775] Step 1:

[0776] Photo and audio data capture

[0777] Users can take their own photos and record audio data using a smartphone or PC, using the camera to capture the photos and the microphone to capture the audio data.

[0778] Input: Smartphone, PC, camera, microphone

[0779] Output: Photo data, audio data

[0780] Step 2:

[0781] Sending photos and audio data

[0782] The user sends the captured photo and audio data to the server via the device, which then packages the data and transfers it to the server via the network.

[0783] Input: Photo data, audio data

[0784] Output: Data transferred to the server

[0785] Step 3:

[0786] Data storage and preprocessing

[0787] The server temporarily stores the received photo and audio data, then analyzes the photo metadata (date and time of shooting, location, etc.) and preprocesses the audio data.

[0788] Input: Data transferred to the server

[0789] Output: Stored photo data, audio data, analyzed metadata

[0790] Step 4:

[0791] Automatic photo processing and compositing

[0792] The server applies AI algorithms to the stored photo data, automatically adjusting brightness and contrast, applying filters, and sometimes even combining multiple photos.

[0793] Input: Saved photo data

[0794] Output: Processed and composited photo data

[0795] Step 5:

[0796] Identifying and displaying relevant photo memories

[0797] The server searches a database to identify past photos relevant to that day and displays these photos along with a special message to the user.

[0798] Input: Parsed metadata

[0799] Output: Display memorable photos, special messages

[0800] Step 6:

[0801] Virtual avatar generation

[0802] The server uses the user's photo and voice data to apply DeepFake technology to generate a digital avatar that resembles the user, which is then used to assist with the try-on experience in the virtual store.

[0803] Input: Saved photo data, audio data

[0804] Output: Digital avatar

[0805] Step 7:

[0806] Try-on support in a virtual store

[0807] The user's digital avatar is used to simulate trying on products in a virtual store, allowing the user to see in real time how the selected products will fit and look.

[0808] Input: Digital avatar, digital product data

[0809] Output: Virtual try-on simulation results

[0810] This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home.

[0811] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0812] The present invention provides an advanced photo editing and AI dialogue system that utilizes user photo and voice data. By combining the system with an emotion engine, the system can recognize user emotions and improve the quality of dialogue and photo editing. The system of the present invention is comprised of a server, a terminal, and a user, with each element playing a different role.

[0813] First, a user takes a photo using a device and sends it to a server. The device provides an interface for sending the photo data and its metadata (such as the date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken during a family trip from the device to the server. This operation causes the server to receive the photos and their metadata.

[0814] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0815] In addition, the server identifies past memorable photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0816] Next, the system incorporates an emotion engine to recognize the user's emotions. The server analyzes the user's voice and facial expression data to identify the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy" and influences the behavior of the entire system.

[0817] The emotion engine dynamically changes the photo processing and filter selection depending on the user's emotional state: for example, if the user is sad, it applies a soothing warm-toned filter, while if the user is happy, it adds vibrant colors and special effects to make the photo even more appealing.

[0818] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, a user can provide their own photo and voice to the server, which then uses DeepFake technology to generate an AI model that looks exactly like the user. An emotion engine also influences this conversation, generating responses based on the user's emotions, making the interaction experience more natural and pleasant.

[0819] This system will enable users to automatically enhance their photos, display memorable photos related to special occasions, and use DeepFake technology and an emotion engine to engage in real-time emotional dialogue with an AI model that resembles them, improving the efficiency of photo management and use and the overall user experience.

[0820] The processing flow will be explained below.

[0821] Step 1:

[0822] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0823] Step 2:

[0824] The device sends the selected photos and metadata to the server, where the actual process of uploading data takes place using Wi-Fi or mobile data.

[0825] Step 3:

[0826] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0827] Step 4:

[0828] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0829] Step 5:

[0830] The server stores the AI-processed photos in a final storage area, along with the original metadata, making them easier to search and view in the future.

[0831] Step 6:

[0832] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0833] Step 7:

[0834] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0835] Step 8:

[0836] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0837] Step 9:

[0838] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0839] Step 10:

[0840] The server uses an emotion engine to analyze the user's emotions, which involves analyzing the user's voice and facial expression data in real time to identify their emotional state. For example, if the user's voice is high-pitched, the emotion engine will recognize "tension."

[0841] Step 11:

[0842] Based on the analysis results of the emotion engine, the server dynamically changes the photo processing and filter selection. For example, if the user is analyzed as feeling depressed, a warm-toned filter that gives a sense of comfort is applied.

[0843] Step 12:

[0844] The user initiates a dialogue with the AI ​​through their device, and the server receives input from the user in real time and generates a response using an AI model based on information from the emotion engine. For example, if the user says, "I'm tired today," the AI ​​will suggest, "Thank you for your hard work. How about this filter to help you relax?"

[0845] Through these steps, the system will be able to automatically edit photos, identify and display memorable photos from the past, and realize real-time dialogue using Deep Fake technology and an emotion engine.

[0846] Example 2

[0847] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0848] Conventional photo editing and AI dialogue systems do not take user emotions into account, resulting in a consistent quality of dialogue and photo editing, limiting the user experience. Furthermore, due to a lack of functionality for effectively utilizing past memorable photos, there was a need for a more comprehensive function for automatically displaying memories related to a specific day. Furthermore, there were insufficient means for realizing natural dialogue in real time based on the user's photos and voice.

[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0850] In this invention, the server includes a means for receiving photos, a means for automatically processing and combining the received photos, a means for identifying and displaying photos associated with memories of that day, a means for engaging in real-time dialogue with an AI generated based on the user's photos and voice, and a means for recognizing the user's emotions and dynamically changing the method of processing the photos according to the emotions. This enables high-quality photo processing and natural dialogue according to the user's emotions, and effectively displays memories associated with a specific day.

[0851] "Means for receiving photos" refers to the function of sending photo data taken by a user via a terminal to a server and receiving that data.

[0852] "Means for automatically processing and combining received photographs" refers to a function that uses an artificial intelligence algorithm to adjust the brightness, contrast, and color tone of received photographs, apply various filters as needed, and combine multiple photographs.

[0853] "Means for identifying and displaying photos that have memories related to that day" refers to the function of searching a database for past photos taken on a specific day and notifying and displaying them to the user.

[0854] "Means for engaging in real-time dialogue with AI generated based on the user's photos and voice" refers to a function that allows users to engage in natural dialogue in real time with a digital agent generated using Deep Fake technology based on the photos and voice data provided by the user.

[0855] "Means of recognizing the user's emotions and dynamically changing the way photos are edited depending on the emotions" refers to a function that uses an emotion engine to analyze the user's voice and facial expression data and dynamically change the way photos are edited depending on their emotional state.

[0856] This invention is a system that enables advanced photo processing and dialogue with AI using photo and voice data. This system is composed of a server, a terminal, and a user, and each element functions in cooperation with the others.

[0857] Hardware and software used

[0858] Server: A high-performance server computer responsible for storing photo and audio data, running AI algorithms, and emotion recognition. It uses cloud storage (e.g., Amazon S3, Google Cloud Storage), AI algorithms (e.g., OpenCV, TensorFlow), emotion engines (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services), and Deep Fake technology (e.g., DeepFaceLab, StyleGAN).

[0859] Device: A smartphone, tablet, computer, etc. that provides an interface for taking photos, recording audio, and transmitting data.

[0860] User: Takes photos, inputs voice, sends and receives data, and interacts with AI.

[0861] Data processing and calculation

[0862] 1. Take and send a photo:

[0863] A user takes a photo using a device and sends it to a server via a dedicated application or browser. For example, a user takes a photo of a scene from a family trip with a smartphone and uploads the photo to a server.

[0864] 2. Photo data storage and processing:

[0865] The server stores the received photos in cloud storage. At the same time, it automatically processes the photos using AI algorithms. Specifically, it uses OpenCV and TensorFlow to adjust brightness, contrast, and color tone, apply filters, and combine multiple photos. For example, the server analyzes a landscape photo and adjusts color tone and contrast to create a clearer image.

[0866] 3. Identify and view photos of past memories:

[0867] The server searches the database for past photos related to a specific date and notifies and displays them to the user. This is done using SQL queries. For example, on a user's birthday, photos taken on past birthdays can be displayed with a special message.

[0868] 4. Emotion recognition and dynamic photo manipulation:

[0869] The server analyzes the user's voice and facial expression data and uses an emotion engine to recognize the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy." The server then dynamically changes the photo editing method and filter selection based on the emotion engine's recognition results. For example, if the user is sad, the server applies a warm-toned filter to edit the photo to match the user's emotion.

[0870] 5. Creating and interacting with AI models using Deep Fake technology:

[0871] The user provides the server with their own photo and voice data. Based on this data, the server uses Deep Fake technology to generate a digital agent that looks exactly like the user. The generated digital agent can then engage in natural conversations with the user in real time. For example, if the user says, "I'm very happy today," the agent will respond, "That's great. Why are you happy?"

[0872] Examples and prompts

[0873] Examples:

[0874] When a user sends photos taken on a family trip from their device to the server, the server automatically processes the photos and displays memories of the trip. The photo processing also changes depending on the user's emotions. Finally, a digital agent generated based on the user's photos and voice data interacts with the user in real time, and the interaction changes depending on the user's emotions.

[0875] Example prompt:

[0876] Submit photos from your family vacation and create an AI model that produces clearer images.

[0877] Search for photos from past birthdays and display them to the user with a special message.

[0878] Analyze a user's photo and voice data, generate a digital agent that resembles the user in real time using Deep Fake technology, and generate responses based on the user's emotions.

[0879] The above is the details regarding the "Description of the Preferred Embodiments."

[0880] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0881] Specific explanation of program processing

[0882] Processing Steps

[0883] Step 1:

[0884] A user takes a photo using a device and sends it to a server. The input data is the photo file and its metadata (date and time of the photo, location, etc.), and the output data is the photo file transferred to the server. In concrete terms, a user takes photos of scenery from a family trip with their smartphone and uploads the photos to the server via a dedicated application.

[0885] Step 2:

[0886] The server temporarily stores the received photo data in cloud storage. The input data is the photo file and metadata sent from the device, and the output data is the photo file stored in cloud storage. Specifically, the server stores the photo data and metadata in Amazon S3 or Google Cloud Storage.

[0887] Step 3:

[0888] The server automatically processes stored photos using AI algorithms. The input data is the stored photo file, and the output data is the processed photo file. Specifically, the server uses OpenCV and TensorFlow to adjust the brightness, contrast, and color tone of the photo and apply filters as needed. For example, it adjusts the color tone and contrast of a landscape photo to create a clearer image.

[0889] Step 4:

[0890] The server identifies past memorable photos related to that day from the database and displays them to the user. The input data is the user's photo database and the current date information, and the output data is the identified past memorable photos. Specifically, the server uses an SQL query to search for past photos taken on a specific day and notifies and displays them to the user. For example, on the user's birthday, photos taken on past birthdays are displayed with a special message.

[0891] Step 5:

[0892] The server analyzes the user's voice and facial expression data and recognizes the user's emotions using an emotion engine. The input data is the user's voice file and facial expression image, and the output data is the recognized emotional state. Specifically, the server analyzes the voice and facial expression using an emotion engine (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services) to identify the user's emotional state. For example, if the user is smiling, it is recognized as "joy."

[0893] Step 6:

[0894] The server dynamically changes the photo processing method according to the emotion. The input data is the recognized emotional state and the photo file, and the output data is the photo file processed according to the emotion. Specifically, the server uses the results of the emotion engine to apply a warm filter to the photo if sadness is recognized, and add a vivid filter or special effect if joy is recognized.

[0895] Step 7:

[0896] The server generates an AI model using DeepFake technology based on the user's photo and voice data. The input data is the user's photo and voice files, and the output data is the generated digital agent. Specifically, the server uses DeepFaceLab and StyleGAN to generate a digital agent that looks exactly like the user.

[0897] Step 8:

[0898] The user interacts with the generated digital agent in real time. The input data is the user's real-time voice and text input, and the output data is the digital agent's response. Specifically, when the user says, "I'm very happy today," the agent responds, "That's great. Why are you happy?", and the interaction progresses naturally.

[0899] The above is a concrete explanation of the processing contents of the program of this system.

[0900] (Application example 2)

[0901] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0902] Existing photo editing systems lack the ability to take user emotions into account and provide AI-based interactive dialogue. In particular, it is difficult to provide personalized responses based on photos and voice, limiting the user experience. Furthermore, the lack of a function to automatically identify and display memorable photos from the past on special occasions leaves users without a way to share their memories more effectively.

[0903] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0904] In this invention, the server includes a means for automatically processing and synthesizing photos, a means for recognizing a user's emotions and improving the quality of photo processing and dialogue in response to the emotions, and a means for generating a digital agent using DeepFake technology based on the user's photo and voice data. This enables photo processing in response to the user's emotions, enabling interactive dialogue that takes emotions into account. Furthermore, by identifying past memorable photos and displaying them on special days, the user's memory sharing experience can be improved.

[0905] The "means for receiving photos" is a function for receiving photo data and its metadata sent from the user's terminal.

[0906] "Means for automatically processing and compositing received photos" refers to a function that uses an AI algorithm to adjust the brightness, contrast, and color tone of received photos, apply filters, and composite multiple photos.

[0907] The "means for identifying and displaying photos with memories related to that day" is a function for identifying past photos related to a date specified by the user and displaying those photos to the user.

[0908] "Means of generating AI based on the user's photo and voice and holding a conversation" refers to a function that uses the user's photo and voice data to generate a digital agent that resembles the user, and that agent then holds a conversation with the user.

[0909] "Means to recognize the user's emotions and improve the quality of photo editing and dialogue according to the emotions" refers to a function that analyzes the user's facial expressions and voice to identify their emotional state, and applies appropriate filters according to that emotional state, improving the quality of dialogue with AI.

[0910] "Means for generating a digital agent resembling a user using DeepFake technology" refers to a function that uses DeepFake technology to create a digital agent resembling a user based on photographs and audio data provided by the user.

[0911] This invention provides an advanced photo processing and AI dialogue system that utilizes user photo and voice data. This system realizes photo reception, emotion recognition, photo processing, specific display of past photos, and interactive dialogue through the application of DeepFake technology.

[0912] Program Description

[0913] Receiving photos

[0914] First, a user takes a photo using their own device and sends it to the server. For this operation, the device provides an interface for sending photo data and its metadata (date and time of shooting, location, etc.) to the server. For example, when a user sends photos taken during a family trip from their device to the server, the server receives the photos and metadata and temporarily stores them.

[0915] emotion recognition

[0916] The server then analyzes the received photos and user-provided audio data, using emotion recognition libraries such as DeepFace and EmotionRecognizer to identify the user's emotional state from their facial expressions and vocal tone. For example, if the user is smiling in the photo, the system will recognize this as "joy."

[0917] Photo editing

[0918] Based on the emotion recognition results, the server automatically processes the photo, including adjusting brightness and contrast, applying filters, and even merging photos. Using AI algorithms, the server applies the optimal processing based on the user's emotional state. For example, if the user is smiling, a vibrant filter is applied, while if not, a warmer color tone is applied.

[0919] Specific display of past photos

[0920] The server also identifies past photo memories associated with the specified date. The feature searches the user's database of photos to find those taken on that date. It notifies the user with a special message, allowing them to share memories more deeply. For example, when a user celebrates their birthday, it presents photos taken on past birthdays.

[0921] Interacting with a digital agent

[0922] Finally, a digital agent is generated using DeepFake technology based on the user's photo and voice data. This allows the user to interact with a digital agent that resembles them in real time. Emotion recognition results are also reflected in the dialogue, and responses are generated that reflect the user's emotions.

[0923] Examples of concrete examples and prompts

[0924] A user takes a photo of a family trip with their smartphone and sends a prompt to the system: "Analyze the emotions of the people in this photo and apply the most appropriate filter." The system recognizes the emotion and returns the photo with a vivid filter applied. Alternatively, by entering a prompt such as "Identify the emotion in this voice and decide which filter to apply to the photo and which story to generate," the system can identify emotions based on the voice data and create appropriate photo editing and story generation.

[0925] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0926] Step 1:

[0927] A user takes a photo on their device and sends the photo and metadata (date, time, location, etc.) from the device to the server. This input includes image data and metadata, which the server receives and temporarily stores. Specifically, the file is uploaded through the device's interface.

[0928] Step 2:

[0929] The server analyzes the received photos and audio data sent by the user. It uses emotion recognition software such as DeepFace or EmotionRecognizer to identify emotions from facial expressions and voice. The input includes image data and audio data, and the output is the emotion recognition results. Specific operations include image processing and audio analysis.

[0930] Step 3:

[0931] The server automatically processes the received photos based on the emotion recognition results. It uses AI algorithms (such as the Python library OpenCV) to adjust brightness, contrast, and color tone, and apply filters. The input includes the original image data and the emotion recognition results, and the output is a processed photo. Specifically, image processing is performed to change the numerical data of the image.

[0932] Step 4:

[0933] The server searches the user's database for photos of past memories related to the specified date and identifies them. The input includes the user's photo database and the specified date, and the output is the identification and extraction of the relevant photos. Specifically, a database search algorithm is executed.

[0934] Step 5:

[0935] The server displays the identified memorable photo to the user. During this process, a special message is attached and notified to the user's device. The input includes the identified memorable photo and the message, and the output is a notification to the user's device. The specific operation is to send the message using the notification system.

[0936] Step 6:

[0937] The server uses DeepFake technology to generate a digital agent based on the user's photo and voice data. The input includes image and voice data, and the generated digital agent is obtained as the output. Specifically, DeepFake modeling and synthesis processing are performed.

[0938] Step 7:

[0939] The user interacts with the generated digital agent in real time. During this process, emotion recognition results are reflected in the dialogue, resulting in more natural responses. The input includes the user's real-time voice data and emotion recognition results, and the digital agent's response is generated as output. Specific operations include natural language processing and real-time speech synthesis.

[0940] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0941] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0942] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0943] [Fourth embodiment]

[0944] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0945] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0946] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0947] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0948] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0949] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0950] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0951] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0952] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0953] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0954] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0955] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0956] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0957] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[0958] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[0959] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[0960] In addition, the server identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. For example, if a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[0961] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, if the user provides their own photo and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to enjoy a conversation with the AI.

[0962] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them using Deep Fake technology, improving the efficiency of photo management and use and enhancing the user experience.

[0963] The processing flow will be explained below.

[0964] Step 1:

[0965] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[0966] Step 2:

[0967] The device sends the selected photos and metadata to the server, where the actual uploading process takes place using Wi-Fi or mobile data.

[0968] Step 3:

[0969] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[0970] Step 4:

[0971] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[0972] Step 5:

[0973] The server stores the AI-processed photos in a final storage area, along with the original metadata, making future searches and display easier.

[0974] Step 6:

[0975] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[0976] Step 7:

[0977] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[0978] Step 8:

[0979] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[0980] Step 9:

[0981] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[0982] Step 10:

[0983] The user initiates a dialogue with the AI ​​through the device, and the server receives input from the user in real time and generates a response using the AI ​​model. For example, if you ask, "What's the weather like today?", the AI ​​will provide weather information in the format desired by the user. This dialogue can be done via text or voice, improving the user experience.

[0984] Through these steps, the system enables automatic photo editing, identification and display of memorable photos, and real-time interaction using Deep Fake technology.

[0985] Example 1

[0986] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0987] In recent years, users have been taking and storing a large number of digital photos, which requires time and effort to organize and edit them. Furthermore, it is difficult to easily look back on past photos of memories, creating a demand for new ways of communicating using photos. Furthermore, while there is hope for the generation and use of AI models that can converse in real time using the user's own photos and voice, achieving this requires advanced technology.

[0988] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0989] In this invention, the server includes means for taking and sending photos from a terminal, means for receiving and storing the photos and their metadata on the server, means for automatically processing and synthesizing the received photos using an AI algorithm, means for searching and displaying past memorable photos related to that day from a database, and means for generating an AI model using DeepFake technology based on the user's photos and voice and conducting real-time conversations with the user. This allows users to automate the organization and processing of photos, easily look back on memorable photos from special days, and enjoy a new communication experience of interacting in real time with an AI model that resembles them.

[0990] A "terminal" is an electronic device that a user uses to take a photo and transmit the captured photo data and its metadata to a server.

[0991] "Photo" refers to image data that is taken by a user using a terminal and sent to a server.

[0992] "Metadata" is auxiliary information that accompanies photo data, such as the date and time the photo was taken and the location where it was taken.

[0993] The "server" is a system that receives photo data and metadata sent from the device, temporarily stores them, and then processes and synthesizes this data using AI algorithms.

[0994] An "AI algorithm" is an artificial intelligence calculation method used by the server to automatically process and synthesize photo data.

[0995] "Automatic processing" is a process that uses AI algorithms to automatically edit received photo data, such as adjusting brightness, contrast, and applying filters.

[0996] "Synthesis" is the process of combining multiple photographic data to generate a new image.

[0997] "Memorable photos" are specific photos taken by the user in the past that are related to that day.

[0998] "Deep Fake technology" is a technology that generates a realistic AI model that resembles a user based on the user's photo and voice data.

[0999] An "AI model" is a digital agent that resembles a user and is generated using Deep Fake technology based on the user's photo and voice data.

[1000] "Real-time conversation" is a form of communication in which a user interacts with a generated AI model and receives an immediate response.

[1001] This invention provides an advanced photo processing and AI dialogue system using user photos and voice data. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[1002] First, a user takes a photo using a device and sends it to a server. A dedicated application is installed on the device, and the user uploads the photo using this application. The device provides a transmission function that includes photo data and its metadata (date and time of photo, location, etc.). A specific example is when a user takes a photo while traveling and sends it to a server via a dedicated application. For example, a user takes a photo of the Eiffel Tower and sends it to a server using the application.

[1003] The server then temporarily stores the received photos and their metadata. The server then automatically processes and combines the stored data using AI algorithms. Specifically, it uses Python scripts and image processing libraries (e.g., OpenCV) to adjust brightness and contrast, apply filters, or combine multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[1004] In addition, the server searches the database for past memory photos related to that day and displays them to the user. This function identifies photos taken on the same day in the past and notifies the user. For example, when a user has a birthday, the server finds photos taken on past birthdays and displays them as "past memories" with a special message.

[1005] Finally, an AI model is generated using DeepFake technology based on the user's photo and voice data, allowing the user to enjoy real-time conversations with a digital agent that resembles them. For example, if a user provides their own face and voice to the server, the server will use DeepFake technology to generate an AI model that looks exactly like the user. This allows the user to have natural conversations with this AI model.

[1006] An example prompt might be:

[1007] "I want to edit a photo of the Eiffel Tower I took during my trip to make it brighter and with more contrast."

[1008] "Find photos from past birthdays and display them with special messages."

[1009] "I would like to use the photos and voice data I provide to generate an AI model that resembles me and then have a conversation with that AI model."

[1010] This system will enable users to automatically edit their own photos, display memorable photos related to special occasions, and even interact with an AI model that resembles them in real time using Deep Fake technology, which is expected to improve the efficiency of photo organization and editing and enhance the user experience.

[1011] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1012] Step 1:

[1013] The user takes a photo on their device, opens a dedicated application, selects the photo, and presses the send button.

[1014] Input: The photo taken and its metadata (date, timestamp, location, etc.).

[1015] Processing: The device uploads the photo data and metadata to the server via a dedicated application.

[1016] Output: Notification of completion of transmission to the server.

[1017] Specific operation: The user takes a photo of a landscape using the camera app on their smartphone, then opens the dedicated app, taps the "Upload Photo" button, and sends the selected photo using the "Send" button.

[1018] Step 2:

[1019] The server receives the photo data and metadata sent from the device and temporarily stores them.

[1020] Input: Photos and metadata sent from the device.

[1021] Processing: The server receives the HTTP request, analyzes the photo data and metadata, and stores the data in a storage service.

[1022] Output: Notification that the photo data and metadata have been saved to storage.

[1023] Specific operation: The server saves the received photos to a storage service (e.g. Amazon S3). After saving is complete, add a record of the save to the database.

[1024] Step 3:

[1025] The server automatically processes and synthesizes the stored photo data using AI algorithms.

[1026] Input: Saved photo data.

[1027] Processing: The server runs a Python script that uses image processing libraries (e.g. OpenCV) to adjust brightness and contrast, apply filters, and possibly blend multiple photos together.

[1028] Output: Processed photo data.

[1029] What it does: The server processes the stored photos sequentially, applying a sepia filter to a photo of the Eiffel Tower, for example, and adjusting the brightness and contrast of another photo.

[1030] Step 4:

[1031] The server searches the database for past memorable photos related to that day, identifies them, and notifies the user.

[1032] Input: Date and time information for that day.

[1033] Processing: The server executes an SQL query to search the database for previous photos taken on the same day. If any are found, the user is notified.

[1034] Output: Identified past memory photos and notification messages.

[1035] What it does: On the user's birthday, the server finds photos taken on past birthdays and notifies the user via the app with a special message.

[1036] Step 5:

[1037] The user sends their own photo and voice data to the server.

[1038] Input: User photo and voice data.

[1039] Processing: The device uploads the photo and audio data to the server via a dedicated application.

[1040] Output: Notification of completion of transmission to the server.

[1041] Specific operation: The user takes a photo of their face and records their voice using a dedicated app, and then sends the data to the server.

[1042] Step 6:

[1043] Based on the received data, the server uses Deep Fake technology to generate an AI model that resembles the user.

[1044] Input: User photo and voice data.

[1045] Processing: The server uses a DeepFake generation tool (e.g., DeepFaceLab) to generate an AI model, which is synthesized based on the user's photo and voice.

[1046] Output: The generated AI model.

[1047] How it works: The server uses deep learning to generate a realistic AI model based on the user's photo and voice data. This AI model is stored on the server and used to interact with the user.

[1048] Step 7:

[1049] Users interact with the generated AI model in real time.

[1050] Input: The generated AI model.

[1051] Processing: The user initiates a conversation with the AI ​​model using a dedicated app. The server runs the necessary infrastructure (e.g., a chat engine) and responds immediately to the user's questions.

[1052] Output: Interaction logs and responses with the AI ​​model.

[1053] How it works: Users communicate with the AI ​​model through text and voice chat. For example, they can ask, "What's the weather like today?" and the AI ​​model will respond in real time.

[1054] Through these steps, users can automatically enhance their photos, easily relive memorable photos from special occasions, and even interact with an AI model that resembles them in real time.

[1055] (Application example 1)

[1056] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1057] Conventional photo editing systems and conversational AI systems have limited advanced interaction and utilization using user photos and voice data, particularly in the virtual store try-on experience. The present invention aims to solve these problems by providing a system that uses user photos and voice data to achieve more advanced photo editing and natural dialogue with a digital agent, and further improves the virtual store try-on experience.

[1058] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1059] In this invention, the server includes means for receiving photos, means for automatically processing and combining the received photos, means for identifying and displaying photos that are memorable for that day, means for generating an AI based on the user's photos and voice to hold a conversation, and means for generating a virtual avatar based on the user's photos and voice to assist in trying on products in a virtual store. This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home, allowing them to check the look and fit of products.

[1060] "Means for receiving photos" means any device or software for sending and receiving photos and their metadata from a user's terminal to the server.

[1061] "Means for automatically processing and synthesizing received photographs" refers to devices or software that use AI algorithms on received photographic data to automatically adjust brightness, contrast, color tone, and apply filters.

[1062] "Means for identifying and displaying memorable photographs relating to that date" refers to a device or software for searching a database for past photographs relating to a particular date and displaying them to a user.

[1063] "Means for generating AI based on a user's photo and voice to converse" refers to a device or software that uses photo and voice data obtained from a user to generate a digital agent that resembles the user using Deep Fake technology, allowing the user to converse naturally with that agent.

[1064] "Means for generating a virtual avatar and assisting in trying on products in a virtual store" refers to a device or software that generates a digital avatar based on a user's photograph and voice, and enables the user to simulate trying on products in a virtual store using that avatar.

[1065] This invention provides an advanced photo processing and AI dialogue system that utilizes photos and voice data taken by users. The system of this invention is composed of multiple elements: a server, a terminal, and a user, each of which plays a specific role.

[1066] The process begins with a user taking a photo using a device and sending the photo to a server. The device provides an interface for sending the photo data and its metadata (date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken while traveling from their device to a server. This operation causes the server to receive the photos and their metadata.

[1067] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms. This includes adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by the user and adjusts color and contrast to create a clearer image. The server also identifies past photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on that day and displays them with a special message.

[1068] Another distinctive feature of this invention is the ability to use DeepFake technology to generate a digital avatar based on the user's photo and voice data, helping them try on products in a virtual store. Users take photos of themselves with their smartphones and provide voice data. This generates a digital avatar, allowing them to try on products in the virtual store. The generated digital avatar allows the user to try on selected products in real time and check their appearance and fit.

[1069] Hardware and software used:

[1070] Smartphone or PC (used for capturing photos and audio)

[1071] Camera (takes photos)

[1072] Microphone (acquires audio data)

[1073] Server (temporary storage of photo and audio data, processing, generation of virtual avatars)

[1074] OpenCV (used for capturing and displaying photos)

[1075] Deep Fake technology (used to generate digital avatars)

[1076] Python script (used to coordinate the entire process)

[1077] Examples:

[1078] Let's say a user wants to try on new clothes. They launch the app and capture their photo and audio data. Based on this, the app generates a digital avatar and shows them how the selected clothes will look in real time. Through virtual try-on, users can see how the products will look at home without having to go to a physical store.

[1079] Example prompt sentence:

[1080] Prompt: "Generate a digital agent that resembles the user based on the following photo and audio data. The photo and audio data are attached. This agent will act as an avatar for the user to try on products in a virtual store."

[1081] This approach allows users to enjoy a virtual store experience and avoids the hassle of visiting a physical store.

[1082] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1083] Step 1:

[1084] Photo and audio data capture

[1085] Users can take their own photos and record audio data using a smartphone or PC, using the camera to capture the photos and the microphone to capture the audio data.

[1086] Input: Smartphone, PC, camera, microphone

[1087] Output: Photo data, audio data

[1088] Step 2:

[1089] Sending photos and audio data

[1090] The user sends the captured photo and audio data to the server via the device, which then packages the data and transfers it to the server via the network.

[1091] Input: Photo data, audio data

[1092] Output: Data transferred to the server

[1093] Step 3:

[1094] Data storage and preprocessing

[1095] The server temporarily stores the received photo and audio data, then analyzes the photo metadata (date and time of shooting, location, etc.) and preprocesses the audio data.

[1096] Input: Data transferred to the server

[1097] Output: Stored photo data, audio data, analyzed metadata

[1098] Step 4:

[1099] Automatic photo processing and compositing

[1100] The server applies AI algorithms to the stored photo data, automatically adjusting brightness and contrast, applying filters, and sometimes even combining multiple photos.

[1101] Input: Saved photo data

[1102] Output: Processed and composited photo data

[1103] Step 5:

[1104] Identifying and displaying relevant photo memories

[1105] The server searches a database to identify past photos relevant to that day and displays these photos along with a special message to the user.

[1106] Input: Parsed metadata

[1107] Output: Display memorable photos, special messages

[1108] Step 6:

[1109] Virtual avatar generation

[1110] The server uses the user's photo and voice data to apply DeepFake technology to generate a digital avatar that resembles the user, which is then used to assist with the try-on experience in the virtual store.

[1111] Input: Saved photo data, audio data

[1112] Output: Digital avatar

[1113] Step 7:

[1114] Try-on support in a virtual store

[1115] The user's digital avatar is used to simulate trying on products in a virtual store, allowing the user to see in real time how the selected products will fit and look.

[1116] Input: Digital avatar, digital product data

[1117] Output: Virtual try-on simulation results

[1118] This allows users to have a realistic try-on experience in a virtual store from the comfort of their own home.

[1119] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1120] The present invention provides an advanced photo editing and AI dialogue system that utilizes user photo and voice data. By combining the system with an emotion engine, the system can recognize user emotions and improve the quality of dialogue and photo editing. The system of the present invention is comprised of a server, a terminal, and a user, with each element playing a different role.

[1121] First, a user takes a photo using a device and sends it to a server. The device provides an interface for sending the photo data and its metadata (such as the date and time of the photo, location, etc.) to the server. A specific example is when a user sends photos taken during a family trip from the device to the server. This operation causes the server to receive the photos and their metadata.

[1122] The server then temporarily stores the received photos and automatically processes and combines them using AI algorithms, including adjusting brightness and contrast, applying filters, or combining multiple photos. For example, the AI ​​analyzes a landscape photo sent by a user and adjusts color and contrast to create a clearer image.

[1123] In addition, the server identifies past memorable photos related to that day and displays them to the user. This function searches the database for photos taken on the same day in the past and notifies the user. Specifically, when a user has a birthday, the server finds photos taken on past birthdays and displays them with a special message.

[1124] Next, the system incorporates an emotion engine to recognize the user's emotions. The server analyzes the user's voice and facial expression data to identify the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy" and influences the behavior of the entire system.

[1125] The emotion engine dynamically changes the photo processing and filter selection depending on the user's emotional state: for example, if the user is sad, it applies a soothing warm-toned filter, while if the user is happy, it adds vibrant colors and special effects to make the photo even more appealing.

[1126] Finally, DeepFake technology is used to generate an AI based on the user's photo and voice data, which then engages in real-time conversations with the user. In this case, the user can engage in natural conversations with a digital agent that resembles them. For example, a user can provide their own photo and voice to the server, which then uses DeepFake technology to generate an AI model that looks exactly like the user. An emotion engine also influences this conversation, generating responses based on the user's emotions, making the interaction experience more natural and pleasant.

[1127] This system will enable users to automatically enhance their photos, display memorable photos related to special occasions, and use DeepFake technology and an emotion engine to engage in real-time emotional dialogue with an AI model that resembles them, improving the efficiency of photo management and use and the overall user experience.

[1128] The processing flow will be explained below.

[1129] Step 1:

[1130] A user takes a photo using a device and selects the photo and its metadata (date and time of the photo, location, etc.).

[1131] Step 2:

[1132] The device sends the selected photos and metadata to the server, where the actual process of uploading data takes place using Wi-Fi or mobile data.

[1133] Step 3:

[1134] The server stores the received photo and metadata in a temporary storage area, after which the server prepares to apply the next processing step to the photo data.

[1135] Step 4:

[1136] The server then passes the stored photos to an AI algorithm, which automatically adjusts the brightness, contrast, and color tone of the photo and applies filters as needed—for example, a filter that enhances natural colors in landscape photos.

[1137] Step 5:

[1138] The server stores the AI-processed photos in a final storage area, along with the original metadata, making them easier to search and view in the future.

[1139] Step 6:

[1140] The server queries the database to identify past photos related to that day. The server takes the current date and searches for photos taken on the same date in the past. For example, if it finds photos taken on a family vacation on the same date in the past, it will include them in the search results.

[1141] Step 7:

[1142] The server prepares to display the found photo of a past memory to the user. Specifically, it links the photo with a message saying "Today's memorable photo" and formats the data for display on the user's device.

[1143] Step 8:

[1144] To initiate a dialogue with the AI, a user provides their own photo and voice data to the server, including a recent photo of their face and a voice recording.

[1145] Step 9:

[1146] Based on the photos and audio data received by the server, Deep Fake technology is used to generate an AI model that resembles the user. The AI ​​model learns the user's characteristics and is ready to engage in natural conversation.

[1147] Step 10:

[1148] The server uses an emotion engine to analyze the user's emotions, which involves analyzing the user's voice and facial expression data in real time to identify their emotional state. For example, if the user's voice is high-pitched, the emotion engine will recognize "tension."

[1149] Step 11:

[1150] Based on the analysis results of the emotion engine, the server dynamically changes the photo processing and filter selection. For example, if the user is analyzed as feeling depressed, a warm-toned filter that gives a sense of comfort is applied.

[1151] Step 12:

[1152] The user initiates a dialogue with the AI ​​through their device, and the server receives input from the user in real time and generates a response using an AI model based on information from the emotion engine. For example, if the user says, "I'm tired today," the AI ​​will suggest, "Thank you for your hard work. How about this filter to help you relax?"

[1153] Through these steps, the system will be able to automatically edit photos, identify and display memorable photos from the past, and realize real-time dialogue using Deep Fake technology and an emotion engine.

[1154] Example 2

[1155] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1156] Conventional photo editing and AI dialogue systems do not take user emotions into account, resulting in a consistent quality of dialogue and photo editing, limiting the user experience. Furthermore, due to a lack of functionality for effectively utilizing past memorable photos, there was a need for a more comprehensive function for automatically displaying memories related to a specific day. Furthermore, there were insufficient means for realizing natural dialogue in real time based on the user's photos and voice.

[1157] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1158] In this invention, the server includes a means for receiving photos, a means for automatically processing and combining the received photos, a means for identifying and displaying photos associated with memories of that day, a means for engaging in real-time dialogue with an AI generated based on the user's photos and voice, and a means for recognizing the user's emotions and dynamically changing the method of processing the photos according to the emotions. This enables high-quality photo processing and natural dialogue according to the user's emotions, and effectively displays memories associated with a specific day.

[1159] "Means for receiving photos" refers to the function of sending photo data taken by a user via a terminal to a server and receiving that data.

[1160] "Means for automatically processing and combining received photographs" refers to a function that uses an artificial intelligence algorithm to adjust the brightness, contrast, and color tone of received photographs, apply various filters as needed, and combine multiple photographs.

[1161] "Means for identifying and displaying photos that have memories related to that day" refers to the function of searching a database for past photos taken on a specific day and notifying and displaying them to the user.

[1162] "Means for engaging in real-time dialogue with AI generated based on the user's photos and voice" refers to a function that allows users to engage in natural dialogue in real time with a digital agent generated using Deep Fake technology based on the photos and voice data provided by the user.

[1163] "Means of recognizing the user's emotions and dynamically changing the way photos are edited depending on the emotions" refers to a function that uses an emotion engine to analyze the user's voice and facial expression data and dynamically change the way photos are edited depending on their emotional state.

[1164] This invention is a system that enables advanced photo processing and dialogue with AI using photo and voice data. This system is composed of a server, a terminal, and a user, and each element functions in cooperation with the others.

[1165] Hardware and software used

[1166] Server: A high-performance server computer responsible for storing photo and audio data, running AI algorithms, and emotion recognition. It uses cloud storage (e.g., Amazon S3, Google Cloud Storage), AI algorithms (e.g., OpenCV, TensorFlow), emotion engines (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services), and Deep Fake technology (e.g., DeepFaceLab, StyleGAN).

[1167] Device: A smartphone, tablet, computer, etc. that provides an interface for taking photos, recording audio, and transmitting data.

[1168] User: Takes photos, inputs voice, sends and receives data, and interacts with AI.

[1169] Data processing and calculation

[1170] 1. Take and send a photo:

[1171] A user takes a photo using a device and sends it to a server via a dedicated application or browser. For example, a user takes a photo of a scene from a family trip with a smartphone and uploads the photo to a server.

[1172] 2. Photo data storage and processing:

[1173] The server stores the received photos in cloud storage. At the same time, it automatically processes the photos using AI algorithms. Specifically, it uses OpenCV and TensorFlow to adjust brightness, contrast, and color tone, apply filters, and combine multiple photos. For example, the server analyzes a landscape photo and adjusts color tone and contrast to create a clearer image.

[1174] 3. Identify and view photos of past memories:

[1175] The server searches the database for past photos related to a specific date and notifies and displays them to the user. This is done using SQL queries. For example, on a user's birthday, photos taken on past birthdays can be displayed with a special message.

[1176] 4. Emotion recognition and dynamic photo manipulation:

[1177] The server analyzes the user's voice and facial expression data and uses an emotion engine to recognize the user's emotional state. For example, if the user is smiling, the emotion engine recognizes this as "joy." The server then dynamically changes the photo editing method and filter selection based on the emotion engine's recognition results. For example, if the user is sad, the server applies a warm-toned filter to edit the photo to match the user's emotion.

[1178] 5. Creating and interacting with AI models using Deep Fake technology:

[1179] The user provides the server with their own photo and voice data. Based on this data, the server uses Deep Fake technology to generate a digital agent that looks exactly like the user. The generated digital agent can then engage in natural conversations with the user in real time. For example, if the user says, "I'm very happy today," the agent will respond, "That's great. Why are you happy?"

[1180] Examples and prompts

[1181] Examples:

[1182] When a user sends photos taken on a family trip from their device to the server, the server automatically processes the photos and displays memories of the trip. The photo processing also changes depending on the user's emotions. Finally, a digital agent generated based on the user's photos and voice data interacts with the user in real time, and the interaction changes depending on the user's emotions.

[1183] Example prompt:

[1184] Submit photos from your family vacation and create an AI model that produces clearer images.

[1185] Search for photos from past birthdays and display them to the user with a special message.

[1186] Analyze a user's photo and voice data, generate a digital agent that resembles the user in real time using Deep Fake technology, and generate responses based on the user's emotions.

[1187] The above is the details regarding the "Description of the Preferred Embodiments."

[1188] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1189] Specific explanation of program processing

[1190] Processing Steps

[1191] Step 1:

[1192] A user takes a photo using a device and sends it to a server. The input data is the photo file and its metadata (date and time of the photo, location, etc.), and the output data is the photo file transferred to the server. In concrete terms, a user takes photos of scenery from a family trip with their smartphone and uploads the photos to the server via a dedicated application.

[1193] Step 2:

[1194] The server temporarily stores the received photo data in cloud storage. The input data is the photo file and metadata sent from the device, and the output data is the photo file stored in cloud storage. Specifically, the server stores the photo data and metadata in Amazon S3 or Google Cloud Storage.

[1195] Step 3:

[1196] The server automatically processes stored photos using AI algorithms. The input data is the stored photo file, and the output data is the processed photo file. Specifically, the server uses OpenCV and TensorFlow to adjust the brightness, contrast, and color tone of the photo and apply filters as needed. For example, it adjusts the color tone and contrast of a landscape photo to create a clearer image.

[1197] Step 4:

[1198] The server identifies past memorable photos related to that day from the database and displays them to the user. The input data is the user's photo database and the current date information, and the output data is the identified past memorable photos. Specifically, the server uses an SQL query to search for past photos taken on a specific day and notifies and displays them to the user. For example, on the user's birthday, photos taken on past birthdays are displayed with a special message.

[1199] Step 5:

[1200] The server analyzes the user's voice and facial expression data and recognizes the user's emotions using an emotion engine. The input data is the user's voice file and facial expression image, and the output data is the recognized emotional state. Specifically, the server analyzes the voice and facial expression using an emotion engine (e.g., Amazon Rekognition, Microsoft Azure Cognitive Services) to identify the user's emotional state. For example, if the user is smiling, it is recognized as "joy."

[1201] Step 6:

[1202] The server dynamically changes the photo processing method according to the emotion. The input data is the recognized emotional state and the photo file, and the output data is the photo file processed according to the emotion. Specifically, the server uses the results of the emotion engine to apply a warm filter to the photo if sadness is recognized, and add a vivid filter or special effect if joy is recognized.

[1203] Step 7:

[1204] The server generates an AI model using DeepFake technology based on the user's photo and voice data. The input data is the user's photo and voice files, and the output data is the generated digital agent. Specifically, the server uses DeepFaceLab and StyleGAN to generate a digital agent that looks exactly like the user.

[1205] Step 8:

[1206] The user interacts with the generated digital agent in real time. The input data is the user's real-time voice and text input, and the output data is the digital agent's response. Specifically, when the user says, "I'm very happy today," the agent responds, "That's great. Why are you happy?", and the interaction progresses naturally.

[1207] The above is a concrete explanation of the processing contents of the program of this system.

[1208] (Application example 2)

[1209] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1210] Existing photo editing systems lack the ability to take user emotions into account and provide AI-based interactive dialogue. In particular, it is difficult to provide personalized responses based on photos and voice, limiting the user experience. Furthermore, the lack of a function to automatically identify and display memorable photos from the past on special occasions leaves users without a way to share their memories more effectively.

[1211] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1212] In this invention, the server includes a means for automatically processing and synthesizing photos, a means for recognizing a user's emotions and improving the quality of photo processing and dialogue in response to the emotions, and a means for generating a digital agent using DeepFake technology based on the user's photo and voice data. This enables photo processing in response to the user's emotions, enabling interactive dialogue that takes emotions into account. Furthermore, by identifying past memorable photos and displaying them on special days, the user's memory sharing experience can be improved.

[1213] The "means for receiving photos" is a function for receiving photo data and its metadata sent from the user's terminal.

[1214] "Means for automatically processing and compositing received photos" refers to a function that uses an AI algorithm to adjust the brightness, contrast, and color tone of received photos, apply filters, and composite multiple photos.

[1215] The "means for identifying and displaying photos with memories related to that day" is a function for identifying past photos related to a date specified by the user and displaying those photos to the user.

[1216] "Means of generating AI based on the user's photo and voice and holding a conversation" refers to a function that uses the user's photo and voice data to generate a digital agent that resembles the user, and that agent then holds a conversation with the user.

[1217] "Means to recognize the user's emotions and improve the quality of photo editing and dialogue according to the emotions" refers to a function that analyzes the user's facial expressions and voice to identify their emotional state, and applies appropriate filters according to that emotional state, improving the quality of dialogue with AI.

[1218] "Means for generating a digital agent resembling a user using DeepFake technology" refers to a function that uses DeepFake technology to create a digital agent resembling a user based on photographs and audio data provided by the user.

[1219] This invention provides an advanced photo processing and AI dialogue system that utilizes user photo and voice data. This system realizes photo reception, emotion recognition, photo processing, specific display of past photos, and interactive dialogue through the application of DeepFake technology.

[1220] Program Description

[1221] Receiving photos

[1222] First, a user takes a photo using their own device and sends it to the server. For this operation, the device provides an interface for sending photo data and its metadata (date and time of shooting, location, etc.) to the server. For example, when a user sends photos taken during a family trip from their device to the server, the server receives the photos and metadata and temporarily stores them.

[1223] emotion recognition

[1224] The server then analyzes the received photos and user-provided audio data, using emotion recognition libraries such as DeepFace and EmotionRecognizer to identify the user's emotional state from their facial expressions and vocal tone. For example, if the user is smiling in the photo, the system will recognize this as "joy."

[1225] Photo editing

[1226] Based on the emotion recognition results, the server automatically processes the photo, including adjusting brightness and contrast, applying filters, and even merging photos. Using AI algorithms, the server applies the optimal processing based on the user's emotional state. For example, if the user is smiling, a vibrant filter is applied, while if not, a warmer color tone is applied.

[1227] Specific display of past photos

[1228] The server also identifies past photo memories associated with the specified date. The feature searches the user's database of photos to find those taken on that date. It notifies the user with a special message, allowing them to share memories more deeply. For example, when a user celebrates their birthday, it presents photos taken on past birthdays.

[1229] Interacting with a digital agent

[1230] Finally, a digital agent is generated using DeepFake technology based on the user's photo and voice data. This allows the user to interact with a digital agent that resembles them in real time. Emotion recognition results are also reflected in the dialogue, and responses are generated that reflect the user's emotions.

[1231] Examples of concrete examples and prompts

[1232] A user takes a photo of a family trip with their smartphone and sends a prompt to the system: "Analyze the emotions of the people in this photo and apply the most appropriate filter." The system recognizes the emotion and returns the photo with a vivid filter applied. Alternatively, by entering a prompt such as "Identify the emotion in this voice and decide which filter to apply to the photo and which story to generate," the system can identify emotions based on the voice data and create appropriate photo editing and story generation.

[1233] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1234] Step 1:

[1235] A user takes a photo on their device and sends the photo and metadata (date, time, location, etc.) from the device to the server. This input includes image data and metadata, which the server receives and temporarily stores. Specifically, the file is uploaded through the device's interface.

[1236] Step 2:

[1237] The server analyzes the received photos and audio data sent by the user. It uses emotion recognition software such as DeepFace or EmotionRecognizer to identify emotions from facial expressions and voice. The input includes image data and audio data, and the output is the emotion recognition results. Specific operations include image processing and audio analysis.

[1238] Step 3:

[1239] The server automatically processes the received photos based on the emotion recognition results. It uses AI algorithms (such as the Python library OpenCV) to adjust brightness, contrast, and color tone, and apply filters. The input includes the original image data and the emotion recognition results, and the output is a processed photo. Specifically, image processing is performed to change the numerical data of the image.

[1240] Step 4:

[1241] The server searches the user's database for photos of past memories related to the specified date and identifies them. The input includes the user's photo database and the specified date, and the output is the identification and extraction of the relevant photos. Specifically, a database search algorithm is executed.

[1242] Step 5:

[1243] The server displays the identified memorable photo to the user. During this process, a special message is attached and notified to the user's device. The input includes the identified memorable photo and the message, and the output is a notification to the user's device. The specific operation is to send the message using the notification system.

[1244] Step 6:

[1245] The server uses DeepFake technology to generate a digital agent based on the user's photo and voice data. The input includes image and voice data, and the generated digital agent is obtained as the output. Specifically, DeepFake modeling and synthesis processing are performed.

[1246] Step 7:

[1247] The user interacts with the generated digital agent in real time. During this process, emotion recognition results are reflected in the dialogue, resulting in more natural responses. The input includes the user's real-time voice data and emotion recognition results, and the digital agent's response is generated as output. Specific operations include natural language processing and real-time speech synthesis.

[1248] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1249] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1250] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1251] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1252] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1253] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1254] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1255] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1256] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1257] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1258] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1259] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1260] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1261] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1262] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1263] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1264] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1265] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1266] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1267] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1268] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1269] The following is further disclosed regarding the above embodiment.

[1270] (Claim 1)

[1271] means for receiving a photograph;

[1272] means for automatically processing and combining the received photographs;

[1273] a means of identifying and displaying memorable photos relating to that day;

[1274] A method to generate AI based on the user's photos and voice and hold a conversation,

[1275] A system including:

[1276] (Claim 2)

[1277] 2. The system of claim 1, wherein the photo receiving means receives the photo and its metadata from the user's terminal.

[1278] (Claim 3)

[1279] 2. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photo using an AI algorithm and applies a filter.

[1280] (Claim 4)

[1281] 2. The system of claim 1, further comprising means for identifying a memorable photograph and searching a database for photographs taken on the same date in the past based on the current date.

[1282] (Claim 5)

[1283] The system of claim 1, wherein the means for generating AI and engaging in conversation uses Deep Fake technology to generate an AI model that looks exactly like the user and uses that AI model to converse with the user.

[1284] "Example 1"

[1285] (Claim 1)

[1286] means for taking and sending photos from the device;

[1287] means for receiving and storing the photos and their metadata on a server;

[1288] A means for automatically processing and synthesizing received photos using AI algorithms;

[1289] A means for searching and displaying past memory photos related to that day from a database;

[1290] A method to generate an AI model using Deep Fake technology based on the user's photo and voice and have a real-time conversation with the user.

[1291] A system including:

[1292] (Claim 2)

[1293] 2. The system according to claim 1, wherein the photo receiving and storing means receives and stores the photo and its metadata sent from the terminal in the server.

[1294] (Claim 3)

[1295] 2. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photo using an AI algorithm and applies a filter.

[1296] "Application Example 1"

[1297] (Claim 1)

[1298] means for receiving a photograph;

[1299] means for automatically processing and combining the received photographs;

[1300] a means of identifying and displaying memorable photos relating to that day;

[1301] A method to generate AI based on the user's photos and voice and hold a conversation,

[1302] A method to generate a virtual avatar based on the user's photo and voice and assist them in trying on products in a virtual store.

[1303] A system including:

[1304] (Claim 2)

[1305] 2. The system of claim 1, wherein the photo receiving means receives the photo and its metadata from the user's terminal.

[1306] (Claim 3)

[1307] 2. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photo using an AI algorithm and applies a filter.

[1308] "Example 2: Combining Emotion Engines"

[1309] (Claim 1)

[1310] means for receiving a photograph;

[1311] means for automatically processing and combining the received photographs;

[1312] a means of identifying and displaying memorable photos relating to that day;

[1313] A means to have real-time conversations with AI generated based on the user's photos and voice, and

[1314] A means to recognize the user's emotions and dynamically change the way the photo is processed depending on the emotions;

[1315] A system including:

[1316] (Claim 2)

[1317] 10. The system of claim 1, wherein the photo receiving means receives the photo and associated metadata from the user's terminal.

[1318] (Claim 3)

[1319] 10. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photograph using artificial intelligence algorithms and applies filters.

[1320] "Application example 2 when combining emotion engines"

[1321] (Claim 1)

[1322] means for receiving a photograph;

[1323] means for automatically processing and combining the received photographs;

[1324] a means of identifying and displaying memorable photos relating to that day;

[1325] A method to generate AI based on the user's photos and voice and hold a conversation,

[1326] A means to recognize the user's emotions and improve the quality of photo editing and interaction according to the emotions;

[1327] A system including:

[1328] (Claim 2)

[1329] 2. The system of claim 1, wherein the photo receiving means receives the photo and its metadata from the user's terminal.

[1330] (Claim 3)

[1331] 2. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photo using an AI algorithm and applies a filter.

[1332] (Claim 4)

[1333] 10. The system of claim 1, wherein the user emotion recognition means analyzes facial and voice data in the photograph to identify an emotional state.

[1334] (Claim 5)

[1335] 5. The system of claim 4, wherein the emotion recognition means is used to apply filters and special effects according to the emotional state.

[1336] (Claim 6)

[1337] The system of claim 1 uses Deep Fake technology to generate a digital agent that resembles a user based on the user's photo and voice data. [Explanation of symbols]

[1338] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving a photograph; means for automatically processing and combining the received photographs; a means of identifying and displaying memorable photos relating to that day; A method to generate AI based on the user's photos and voice and hold a conversation, A system including:

2. 2. The system of claim 1, wherein the photo receiving means receives the photo and its metadata from a user terminal.

3. 2. The system of claim 1, wherein the automatic processing and compositing means adjusts the brightness, contrast, and color tone of the photo using AI algorithms and applies filters.

4. 2. The system of claim 1, further comprising means for identifying a memorable photograph and searching a database for photographs taken on the same date in the past based on the current date.

5. The system of claim 1, wherein the means for generating AI and engaging in conversation uses Deep Fake technology to generate an AI model that looks exactly like the user and uses that AI model to converse with the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A