system

A system for dual-income households automates the selection, editing, and sharing of children's photos and videos, addressing organization challenges and enabling targeted advertising.

JP2026036341APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Dual-income households face challenges in organizing and sharing photos and videos of their children's growth efficiently, as they often contain unwanted backgrounds and reflections, and traditional methods lack effective ways to automate this process and target advertisements based on recipient attributes.

Method used

A system that includes uploading and storing photos and videos, selecting high-quality content based on evaluation criteria, automatically correcting backgrounds and reflections, organizing into albums, and sharing with designated recipients, while also generating targeted advertisements.

Benefits of technology

The system automates the organization and sharing of high-quality photos and videos, reducing user workload and enabling accurate advertising based on recipient attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036341000001_ABST
    Figure 2026036341000001_ABST
Patent Text Reader

Abstract

Provide a system. A means for uploading photos and videos taken; A means to retrieve and store metadata for uploaded photos and videos; A means for selecting photos and videos within a specified period based on evaluation criteria; A means to automatically correct unwanted backgrounds and reflections in selected photos and videos, Edited photos and videos can be organized into albums and automatically shared with designated recipients. A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, the number of dual-income households has increased, and parents tend to take many photos and videos to record their children's growth. However, there is a challenge in organizing these photos and videos in their busy daily lives and regularly sharing them with family members such as grandparents. Furthermore, the photos and videos taken often contain unnecessary backgrounds and reflections, which can make them look unattractive. To solve these issues, a system that automates the process of selecting, editing, and sharing photos and videos is needed. [Means for solving the problem]

[0005] The present invention is a system that includes a means for uploading photographs and videos and acquiring and storing their metadata. It also includes a means for selecting photographs and videos from a specified period based on evaluation criteria such as facial expression, composition, lighting, and the presence or absence of blur and noise. It also includes a means for automatically correcting unwanted backgrounds and reflections in the selected photographs and videos. It also includes a means for compiling the corrected photographs and videos into an album and automatically sharing them with designated recipients (e.g., grandparents). It also includes a means for generating advertising targeting data based on recipient attributes. This system allows dual-income families to efficiently organize and share records of their children's growth.

[0006] "Photographs and videos" refer to digital data of still images and videos taken by the user.

[0007] "Means for uploading" refers to the function for transferring digital data from a user's terminal to a server.

[0008] "Metadata" refers to additional information associated with photos and videos, such as the date and time they were taken, location, and device information.

[0009] "Capture and preservation means" refers to the functionality for extracting metadata from uploaded digital data and recording it in a database or cloud storage.

[0010] "Specified period" refers to a specific time frame that the system is set to by a user or administrator.

[0011] "Evaluation criteria" are the criteria that the AI ​​algorithm uses to select photos and videos, including facial expressions, composition, lighting, and the presence or absence of blur or noise.

[0012] "Means of selection" refers to the function of using an AI algorithm to select photos and videos based on evaluation criteria.

[0013] "Correction methods" refers to the ability to use AI technology to detect unwanted backgrounds or reflections in photos and videos and remove or replace them.

[0014] "Album organization" refers to the ability to organize selected and modified photos and videos into a user-friendly layout and display them in a specific format.

[0015] A "designated recipient" refers to a person who has an email address or application account that the user has registered in advance.

[0016] "Automatic sharing means" refers to the ability to notify or forward the modified and organized album to specific recipients.

[0017] "Recipient attributes" refers to information that indicates the interests and behavior of users viewing photos and videos, the frequency of taking photos, the content of photos, etc.

[0018] "Advertising targeting data" means data used to target specific advertising content based on recipient attributes. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken with multiple functions in a server that processes that data.

[0041] 1. Uploading photos and videos

[0042] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0043] 2. Data storage and initial processing

[0044] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0045] 3. AI-powered photo and video selection

[0046] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. This process extracts only high-quality photos and videos from a huge number of photos and videos.

[0047] 4. Automatic correction of background and unwanted reflections

[0048] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0049] 5. Create and share albums

[0050] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0051] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​selects the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0052] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, a user with many travel photos can be served travel-related ads, while a user with many photos of sporting events can be served sporting goods ads. This allows for highly accurate targeting.

[0053] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] A user takes a photo or video using a smartphone or camera.

[0057] Step 2:

[0058] The user opens a dedicated application and selects the photos and videos they have taken.

[0059] Step 3:

[0060] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0061] Step 4:

[0062] The server receives the uploaded photos and videos.

[0063] Step 5:

[0064] The server stores the received photos and videos in cloud storage.

[0065] Step 6:

[0066] The server extracts metadata from photos and videos (date and time of shooting, location, device information) and stores it in a database.

[0067] Step 7:

[0068] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0069] Step 8:

[0070] An AI algorithm on the server analyzes the photos and videos in the list based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects the best photos and videos.

[0071] Step 9:

[0072] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0073] Step 10:

[0074] The server then saves the edited photos and videos back to cloud storage.

[0075] Step 11:

[0076] The server organizes the selected and edited photos and videos into an album.

[0077] Step 12:

[0078] The server generates an album and sends a sharing notification to the specified recipient's email address or dedicated application.

[0079] Step 13:

[0080] The recipient will receive a link or notification to view the album.

[0081] Step 14:

[0082] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0083] Step 15:

[0084] The server generates targeted advertising data based on user attributes.

[0085] Step 16:

[0086] The server provides the ad targeting data to the relevant companies.

[0087] Example 1

[0088] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0089] In today's dual-income households, it is difficult to efficiently organize and share records of children's growth with family members. To solve this problem, a system that allows users to easily upload, organize, select, edit, and share photos and videos is needed. It also needs the ability to automatically select only high-quality photos and videos from a large number of images and correct backgrounds and unwanted reflections. Furthermore, it is also important to target advertisements based on the recipient's attributes using the shared records.

[0090] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0091] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for saving selected photos and videos in cloud storage and saving the metadata in a database, a means for analyzing and selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos based on the evaluation criteria, a means for compiling the corrected photos and videos into an album and automatically sharing them with designated recipients, a means for sending a sharing link for the album to the recipients via email, and a means for sending a notification to a dedicated application. This allows dual-income households to efficiently organize their children's growth records and easily share them with their families. It also enables advertising targeting based on recipient attributes.

[0092] "Means for uploading photos and videos" refers to a hardware or software interface that provides the functionality for users to send photos and videos they have taken to a server.

[0093] "Means for acquiring and storing metadata" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from uploaded photos and videos, and stores it in a database.

[0094] "Means for saving to cloud storage" refers to the ability to store uploaded photos and videos on a remote server via the Internet, making them accessible.

[0095] The "means of analyzing and selecting based on evaluation criteria" is a function that evaluates photos and videos taken within a specified period of time using criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise, and selects high-quality photos and videos.

[0096] "Automatic correction" is a function that detects unwanted backgrounds and reflections in selected photos and videos and automatically processes them to improve their appearance.

[0097] "A means of organizing photos and videos into an album and automatically sharing them with designated recipients" is a function that lays out edited photos and videos in an album format and shares the album by sending a link or notification to designated recipients.

[0098] The "means for sending a shared link of an album by email" is a function for sending a link of a created album to a recipient's email address.

[0099] The "means for sending a notification to a dedicated application" is a function for sending a notification about the created album to a recipient as a push notification via a dedicated application.

[0100] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system significantly reduces the user's workload by automating the process of uploading, saving, selecting, editing, and sharing photos and videos.

[0101] Uploading photos and videos

[0102] Users upload photos and videos they have taken to the server using a dedicated application (for example, the GrowMemories app). This application provides users with a simple interface and has the ability to select and upload multiple files at once. When uploading, the device compresses the selected photos and videos and sends them to the server.

[0103] Data storage and metadata extraction

[0104] When the server receives uploaded photos and videos, it stores them in cloud storage (e.g., AWS (registered trademark) S3). Once storage is complete, the server extracts metadata (such as the date and time of the photo, location, and device information). This metadata is obtained using a dedicated metadata extraction library (e.g., ExifTool). The extracted metadata is stored in a database (e.g., Amazon RDS). If the file format is inappropriate, the server generates an error message and notifies the user.

[0105] AI-powered photo and video selection

[0106] The AI ​​algorithm (e.g., Google® Cloud Vision API) on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (e.g., facial expression, composition, lighting, presence or absence of blur and noise), allowing it to select the highest quality photos and videos.

[0107] Auto-correct photos and videos

[0108] The server uses generative AI (for example, Adobe Photoshop API) to automatically correct backgrounds and unwanted reflections in the selected photos and videos. Automatic correction removes other people and unwanted objects in the background, resulting in better-looking photos and videos.

[0109] Create and share albums

[0110] The server organizes the edited photos and videos into albums, which are automatically generated in HTML or PDF format. The albums are then shared with designated recipients (e.g., grandparents) via a link sent to their email addresses (e.g., using SendGrid) or a notification sent to a dedicated application (e.g., using Firebase Cloud Messaging).

[0111] Specific examples

[0112] For example, consider the case where Mr. A and his family, a dual-income household, use this system. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. After taking the photos, he uploads the data through a dedicated application (the GrowMemories app). The server saves the data in cloud storage, extracts metadata, and stores it in a database. An AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server uses the Adobe Photoshop API to correct unnecessary parts of the background. The server then creates an album based on the corrected photos and videos and sends a link to Mr. A's parents (the child's grandparents) by email. As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow.

[0113] An example of a prompt sentence when using a generative AI model can be written as follows:

[0114] Example prompt:

[0115] "Please explain the programming process of a system that allows dual-income families to efficiently organize and share records of their children's growth. This system has a mechanism whereby users upload photos and videos they have taken, and the data is stored and processed on a server. Please explain in detail how the data will be organized and shared, including the names of the specific hardware and software."

[0116] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0117] Step 1: Upload photos and videos

[0118] The user selects the photos and videos they have taken using a dedicated application. Multiple files can be selected at the same time using the bulk selection function. When the user presses the upload button, the device compresses the selected files and sends them to the server. The input here is the photo or video file they have taken, and the output is the file sent to the server.

[0119] Step 2: Storing data and extracting metadata

[0120] When the server receives uploaded files, it first saves them to cloud storage. It receives uploaded photo and video files as input and saves them to cloud storage as output. After saving is complete, the server extracts metadata for each file. It uses a metadata extraction library to obtain the shooting date and time, location, device information, etc. The input is the saved photo or video file, and the output is the extracted metadata.

[0121] Step 3: Storing Metadata and Error Notification

[0122] The server saves the extracted metadata in a database. If the file format is inappropriate, it sends an error message to the user based on that information. The input is the extracted metadata and the file format check results, and the output is saving the metadata in the database and sending an error message.

[0123] Step 4: AI-powered photo and video selection

[0124] An AI algorithm on the server analyzes photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur and noise. The input is photo and video files from the specified period and their metadata, and the output is a selection of high-quality photos and videos.

[0125] Step 5: Run Auto Fix

[0126] The server uses generative AI to automatically correct backgrounds and unwanted reflections from selected photos and videos. Corrections include removing other people and unnecessary objects. The input is the selected photo or video file, and the output is the corrected photo or video.

[0127] Step 6: Create an album and share it

[0128] The server uses the corrected photos and videos to compile them into an album and generate it in HTML or PDF format. The input is the corrected photos and videos, and the output is the generated album. To share this album with specified recipients, a link is sent to the recipient's email address or a notification is sent to a dedicated application. An email sending library is used to send emails, and a push notification service is used for notifications. The input is the generated album and recipient information, and the output is the shared album.

[0129] Specific examples of operation

[0130] For example, consider the case where Mr. A and his family, a dual-income household, use this system. They go to the zoo on their day off and take 20 photos and 10 videos. After taking the photos, they upload the data to the server using a dedicated application. The server stores the data in cloud storage, extracts metadata, and stores it in a database. Next, an AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server then uses generative AI to correct unwanted background parts. The server then compiles the corrected photos and videos into an album and emails a link to Mr. A's parents. As a result, Mr. A can easily share organized records, and the grandparents can enjoy watching their grandchildren grow up.

[0131] (Application example 1)

[0132] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0133] Today's busy dual-income households make it difficult to organize photos and videos and regularly share them with grandparents and other family members. It's also a tedious task to select high-quality photos and videos from the many available and remove unwanted backgrounds and unwanted reflections. Furthermore, traditional methods for delivering ads tailored to users' interests have limited accuracy, making effective targeting difficult.

[0134] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0135] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients, a means for analyzing specified media data and estimating user interests and concerns, and a means for targeting and delivering advertisements based on the estimation results. This allows dual-income families to efficiently organize and share records of their children's growth, and enables companies to effectively target advertisements.

[0136] "Means for uploading photos and videos" refers to a function that provides an interface for users to transfer photos and videos taken with their smartphones or cameras to a server via the Internet.

[0137] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts metadata such as the date and time of shooting, location, and device information from uploaded media data and stores it in a database.

[0138] "A means of selecting photos and videos from a specified period based on evaluation criteria" is a function that analyzes and selects photos and videos uploaded within a specific period based on criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[0139] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses an AI algorithm to automatically remove or correct unwanted backgrounds and reflections in selected photos and videos.

[0140] "A means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that organizes edited photos and videos into albums and automatically shares them with designated recipients via links or notifications.

[0141] "Means of analyzing specified media data and inferring user interests and concerns" refers to a function that uses an AI algorithm to analyze the content of uploaded photos and videos and predict user interests and concerns.

[0142] "Means for targeting and delivering advertisements based on inferred results" refers to a function that delivers advertisements optimized for individual users based on user profile information obtained from content analysis of photos and videos.

[0143] The system for implementing this invention allows users to upload photos and videos they have taken and process the data on a server to realize advertising targeting based on their interests. The specific configuration and operating procedures of the system are described below.

[0144] System Configuration

[0145] The system mainly includes the following elements:

[0146] 1. User devices: smartphones and digital cameras

[0147] 2. Servers: Cloud storage, AI algorithms, database servers

[0148] 3. Communication Infrastructure: Internet

[0149] Uploading photos and videos

[0150] Users upload photos and videos taken with their smartphones or digital cameras to a server via a dedicated application. The dedicated application provides an interface that allows users to select multiple files at once and easily upload them. The cloud storage used here could be Google Cloud Storage or Amazon S3.

[0151] Retrieving and storing metadata

[0152] The server extracts metadata such as the date and time of the photo or video, location, and device information from the uploaded photo or video, and stores it in a database. The database typically used is PostgreSQL or MySQL (registered trademark).

[0153] Selection of photos and videos within a specified period

[0154] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur or noise.

[0155] Automatic correction of unwanted backgrounds and reflections

[0156] For selected photos and videos, a generative AI model on the server automatically corrects unnecessary backgrounds and reflected elements.

[0157] Create and share albums

[0158] The edited photos and videos are organized into an album, and the server automatically shares them with designated recipients in the form of a link or notification.

[0159] User interest estimation and ad targeting

[0160] The server analyzes the uploaded media data using an AI algorithm to estimate the user's interests and preferences, and then delivers targeted advertisements optimized for the user based on these estimates.

[0161] Specific examples

[0162] For example, a user might upload a photo of a beach they took on vacation. The photo is stored on a server, and AI analyzes the beach image and classifies it as "travel." The server then generates a travel-related ad (e.g., "Special discounts to exotic destinations!") and displays it in the user's app.

[0163] Prompt Sentence Examples

[0164] Here's an example of a prompt to feed to a generative AI model:

[0165] Based on the results of photo analysis, predict user interests and generate appropriate ads.

[0166] In this way, dual-income families can efficiently organize and share records of their children's growth, and companies can perform highly accurate advertising targeting.

[0167] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0168] Step 1:

[0169] Taking and uploading photos and videos

[0170] A user takes photos or videos with a smartphone or digital camera. The captured media data is uploaded to a server via a dedicated application. The input is the captured photos or videos, and the output is the media data stored in cloud storage.

[0171] Step 2:

[0172] Retrieving and storing metadata

[0173] The server extracts metadata (such as the date and time of the photo, location, and device information) from the uploaded photos and videos and stores it in a database. The input is the uploaded photos and videos, and the output is the metadata stored in the database. Specifically, the system uses a software module to extract the metadata.

[0174] Step 3:

[0175] Photo and video selection

[0176] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise). The input is the saved photos and videos, and the output is a selected set of high-quality photos and videos.

[0177] Step 4:

[0178] Correcting unwanted backgrounds and reflections

[0179] The generative AI model in the server automatically corrects unwanted backgrounds and reflected elements in the selected photos and videos. The input is the selected photos and videos, and the output is the corrected photos and videos. The specific operation is to apply a background removal algorithm.

[0180] Step 5:

[0181] Create and share albums

[0182] The server automatically shares the edited photos and videos with the designated recipients by organizing them into albums. The input is the edited photos and videos, and the output is the album link and notification that the recipients can access. Specifically, the album generation software is used.

[0183] Step 6:

[0184] User interest estimation

[0185] The server analyzes the content of the uploaded media data and infers the user's interests. The input is the content information of the uploaded photos and videos, and the output is the inferred results regarding the user's interests. The specific operation is to apply an AI analysis model.

[0186] Step 7:

[0187] Ad Targeting and Delivery

[0188] Advertisements are targeted and delivered based on the inference results. The input is the inference results regarding the user's interests, and the output is targeted advertisements. Specifically, the ad generation module is used to select and deliver appropriate advertisements.

[0189] In this way, by taking specific actions at each step, dual-income families can efficiently organize and share their children's growth records, and companies can achieve highly accurate advertising targeting.

[0190] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0191] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine.

[0192] 1. Uploading photos and videos

[0193] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0194] 2. Data storage and initial processing

[0195] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0196] 3. Photo and video selection using AI and emotion engine

[0197] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. In addition, an emotion engine recognizes the emotions of the subjects in the photos and videos based on facial recognition, and makes selections based on that emotional information. This process extracts high-quality photos and videos that express rich emotions.

[0198] 4. Automatic correction of background and unwanted reflections

[0199] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0200] 5. Create and share albums

[0201] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0202] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0203] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0204] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0205] The processing flow will be explained below.

[0206] Step 1:

[0207] A user takes a photo or video using a smartphone or camera.

[0208] Step 2:

[0209] The user opens the dedicated application and selects the photos and videos they have taken.

[0210] Step 3:

[0211] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0212] Step 4:

[0213] The server receives the uploaded photos and videos and stores them in cloud storage.

[0214] Step 5:

[0215] The server extracts metadata (such as the date and time of the photo, location, and device information) from the photos and videos it receives and stores it in a database.

[0216] Step 6:

[0217] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0218] Step 7:

[0219] An AI algorithm on the server analyzes photos and videos based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects "good" photos and videos.

[0220] Step 8:

[0221] The emotion engine in the server recognizes the faces of the subjects in the selected photos and videos and recognizes their emotions (joy, sadness, surprise, etc.).

[0222] Step 9:

[0223] Based on the emotional information recognized by the emotion engine, the server prioritizes the selection of high-quality photos and videos that express emotions richly.

[0224] Step 10:

[0225] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0226] Step 11:

[0227] The server then saves the edited photos and videos back to cloud storage.

[0228] Step 12:

[0229] The server organizes the edited photos and videos into an album.

[0230] Step 13:

[0231] The server automatically shares the generated album with the designated recipients (e.g. grandparents) by sending a link to their email address or by sending a notification to their application.

[0232] Step 14:

[0233] The recipient will receive a link or notification to view the album.

[0234] Step 15:

[0235] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0236] Step 16:

[0237] The server generates targeted advertising data based on user attributes, and also reflects emotional information recognized by the emotion engine in the targeting data.

[0238] Step 17:

[0239] The server provides the ad targeting data to the relevant companies.

[0240] Example 2

[0241] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0242] It has been difficult for dual-income families to efficiently organize records of their children's growth and regularly share them with grandparents and other family members. Also, sorting and selecting photos and videos taken, and correcting unnecessary backgrounds and reflections, was time-consuming and laborious. Furthermore, it was difficult to select emotionally rich photos and videos.

[0243] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0244] In this invention, the server includes a means for uploading captured photos and videos, a means for storing the uploaded photos and videos in cloud storage and acquiring and storing metadata in a database, a means for selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and an emotion engine, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos using generative AI, and a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients via email or notification. This makes it possible to organize a child's growth record in a high-quality format and easily share it with family members.

[0245] "Means for uploading photos and videos taken" is a function for sending data of photos and videos taken by the user to a server.

[0246] "Means of storing uploaded photos and videos in cloud storage, obtaining metadata, and storing it in a database" refers to a function that stores photos and videos received by the server on the cloud, extracts information (metadata) such as the date and time the file was taken and the location, and stores it in a database.

[0247] "A means of selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and emotion engine" is a function that uses AI to analyze and evaluate photos and videos within a specified period based on criteria such as facial expressions and composition, and then analyzes their emotional state using an emotion engine to select them.

[0248] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos using generative AI" refers to a function that uses AI technology to detect unwanted parts of the background or other objects contained in selected photos and videos, and automatically corrects or removes them.

[0249] "A means of organizing edited photos and videos into an album and automatically sharing them with designated recipients via email or notification" refers to a function that organizes edited data into an album and sends a link to the recipient's email address or automatically shares it via notification.

[0250] This invention is an information processing system that allows dual-income families to efficiently organize their children's growth records and regularly share them with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine and generative AI.

[0251] First, users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once. During uploading, data is transferred securely using SSL / TLS.

[0252] The server then stores the uploaded photos and videos in cloud storage such as Amazon S3. At the same time, it extracts metadata from the photos and videos, such as the date and time of the photo, location, and device information, and stores the metadata in a database (e.g., MySQL). It checks the integrity of the data and notifies the user if the file format is inappropriate.

[0253] Then, an AI algorithm (e.g., the TENSORFLOW® model) stored on the server evaluates photos and videos from a specified period (e.g., the past month). Evaluation criteria include facial expression, composition, lighting, and the presence or absence of blur and noise. An emotion engine is also used to select photos and videos rich in emotion. This allows high-quality photos and videos that express rich expressions and emotions to be identified.

[0254] Next, a generative AI (e.g., a GAN model) on the server automatically detects and corrects background and unwanted reflections in the selected photos and videos. For example, it removes other people or unnecessary objects in the background to create a more attractive photo or video.

[0255] Finally, the server automatically organizes the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0256] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A goes to the zoo with his family on a day off and takes photos and videos. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0257] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0258] An example of a prompt might be:

[0259] "Create an album of your best family photos from the past month, removing unwanted backgrounds."

[0260] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0261] Step 1:

[0262] Users select photos and videos they have taken using a dedicated application and upload them to the server through a dedicated interface.

[0263] Input: Photos and video files taken by the user

[0264] Output: Photo and video files uploaded to the server

[0265] Specific example of operation: A user selects all photos and videos of the zoo taken with their smartphone in a dedicated app and clicks the upload button.

[0266] Step 2:

[0267] The server stores the uploaded photos and videos in cloud storage (e.g., Amazon S3), extracts metadata from each file (e.g., date and time of capture, location, device information), and stores it in a database (e.g., MySQL).

[0268] Input: Uploaded photo and video files

[0269] Output: Photos and videos stored in cloud storage, metadata stored in a database

[0270] Specific example of operation: The server saves the received files in Amazon S3 and stores the date, time, and location of the photo in a database.

[0271] Step 3:

[0272] An AI algorithm (e.g., TensorFlow model) on the server analyzes photos and videos taken within a specified period based on evaluation criteria (e.g., facial expression, composition, lighting, presence or absence of blur and noise). The emotion engine analyzes emotions using facial recognition and selects high-quality photos and videos based on this information.

[0273] Input: Photo and video files retrieved from cloud storage, metadata from a database

[0274] Output: High-quality photo and video files selected based on evaluation criteria

[0275] Specific example of operation: The AI ​​model analyzes photos from the past month and selects the family photo with the most smiling faces.

[0276] Step 4:

[0277] Generative AI (e.g., GAN model) within the server automatically detects and corrects unwanted backgrounds and reflections in selected photos and videos.

[0278] Input: Selected photo and video files

[0279] Output: Photo or video file with unwanted background and reflections removed

[0280] Specific example of operation: Removes other people in the background and processes the photo into a clean one.

[0281] Step 5:

[0282] The server organizes the modified photos and videos into an album, and generates links and notifications to automatically share the album with the specified recipients.

[0283] Input: Modified photo or video files

[0284] Output: Digital content in album format, sharing links and notifications

[0285] How it works: The server creates a digital album with the modified family photos and emails the link to the grandparents.

[0286] (Application example 2)

[0287] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0288] In dual-income households, efficiently organizing and sharing the large number of photos and videos taken in a high-quality format is a significant burden. There is also a need to reduce the workload involved in sharing children's growth records with family members, and for companies to effectively target advertisements based on user data. Given this background, a system that easily organizes, edits, and automatically shares photos and videos is needed. Furthermore, a means is needed to efficiently deliver personalized advertisements using stored data and emotional information.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0290] In this invention, the server includes means for uploading captured photos and videos, means for acquiring and storing metadata for the uploaded photos and videos, means for selecting photos and videos within a specified period based on evaluation criteria, means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, means for organizing the corrected photos and videos into an album format and automatically sharing them with specified recipients, and means for generating advertising targets based on the stored data and user profiles and displaying personalized advertisements using emotional data. This allows dual-income households to efficiently organize and share photos and videos, and enables companies to easily deliver targeted advertisements based on users' interests.

[0291] - "Means for uploading photos and videos taken" refers to a function that allows users to send photos and videos they have taken to a server via the Internet.

[0292] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from photos and videos received by the server and stores it in a database.

[0293] "A means of selecting photos and videos taken within a specified period based on evaluation criteria" is a function that analyzes photos and videos taken within a specific period based on evaluation criteria such as image quality, composition, facial expression, and lighting conditions, and selects the most suitable ones.

[0294] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses AI to automatically correct unnecessary elements contained in selected photos and videos, improving their quality.

[0295] "Means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that automatically compiles edited photos and videos into albums and electronically shares them with specified recipients.

[0296] "Means for generating advertising targets based on stored data and user profiles, and displaying personalized advertisements using emotional data" refers to a function that generates optimal advertising targets based on the user's stored data, profile information, and emotional monitoring data, and displays them to the user.

[0297] To implement this invention, we need to build a system that efficiently manages and shares photos and videos taken by users. This system consists of major components including smartphones, cloud storage, AI algorithms, emotion engines, and ad distribution platforms.

[0298] System program and processing flow

[0299] 1. Data upload

[0300] Users use a smartphone to upload photos and videos taken using a dedicated application to the server. The uploaded data is stored in cloud storage. During this process, users are provided with a simple interface and can upload multiple files at once.

[0301] 2. Metadata Acquisition and Storage

[0302] Metadata (e.g., shooting date and time, location, device information) is extracted from photos and videos stored in cloud storage and stored in a database. Image processing software (e.g., Google Cloud Vision) is used to obtain the metadata.

[0303] 3. Data selection using AI

[0304] An AI algorithm (e.g., OpenAI® GPT-4®) analyzes photos and videos from a specified period based on evaluation criteria, such as image quality, composition, facial expression, lighting, and the presence or absence of blur and noise, to select high-quality data.

[0305] 4. Emotion Recognition and Data Modification

[0306] An emotion engine (e.g., Microsoft®'s Azure® Emotion API) recognizes facial expressions in photos and videos and extracts emotional data. This emotional data is then reflected in the photo and video selection process. Furthermore, generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0307] 5. Album creation and sharing

[0308] The server automatically compiles the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0309] 6. Ad Targeting and Delivery

[0310] Generate advertising targets based on stored data, user profiles, and emotional data. Display personalized ads to the generated advertising targets. This is done using an advertising distribution platform (e.g., Google AdSense).

[0311] Examples of concrete examples and prompts

[0312] As a concrete example, a dual-income household user uploads photos and videos taken at a park on a day off through a dedicated application. The server stores these data in cloud storage, retrieves metadata, and organizes them. AI and an emotion engine select the best photos of the month and correct unwanted background parts. The system then creates an album using the corrected photos and videos, and automatically shares the link with grandparents. Based on the user's profile and emotion data, personalized travel-related advertisements are displayed.

[0313] Examples of prompts:

[0314] Analyze the metadata of photos and videos uploaded by user A over the past month, and use the emotional data obtained from the emotion engine to generate the most appropriate advertising targets.

[0315] Based on the user's emotional data and activity history, create and deliver ads that are likely to be of interest to them. For example, if there are a lot of travel photos and videos, generate travel-related ads, or if there are a lot of photos of sporting events, generate ads for sports equipment.

[0316] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0317] Step 1:

[0318] Users use their smartphones to upload the photos and videos they have taken to the server via a dedicated application.

[0319] Input: Photo and video files taken with a smartphone.

[0320] Output: Photo and video files saved in cloud storage.

[0321] What it does: A user opens the app, selects a photo or video to upload, and the app makes a request to send the file to the server, which receives it and stores it in cloud storage (e.g., AWS S3, Google Cloud Storage).

[0322] Step 2:

[0323] The server extracts metadata (e.g., shooting date and time, location, device information) from photos and videos stored in cloud storage and stores it in a database.

[0324] Input: Photo and video files stored in cloud storage.

[0325] Output: Metadata stored in a database.

[0326] What it does: The server uses image processing software (e.g., Google Cloud Vision) to analyze and extract metadata from photos and videos and store it in a database.

[0327] Step 3:

[0328] The server analyzes photos and videos from a specified period based on evaluation criteria and selects high-quality data.

[0329] Input: Metadata stored in the database and photo and video files.

[0330] Output: A list of selected high-quality photos and videos.

[0331] How it works: The server uses an AI algorithm (e.g., OpenAI GPT-4) to analyze photos and videos from a specified period (e.g., the past month) based on evaluation criteria (image quality, composition, facial expression, lighting conditions, presence or absence of blur and noise), and selects the most suitable ones.

[0332] Step 4:

[0333] The server uses an emotion engine to obtain emotion data from selected photos and videos and corrects the data.

[0334] Input: Selected photos and videos, and user profile information.

[0335] Output: The retouched photo or video files.

[0336] How it works: An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in selected photos and videos to obtain emotional data. Then, a generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0337] Step 5:

[0338] The server organizes the edited photos and videos into an album and automatically shares them with the designated recipients.

[0339] Input: The modified photo or video file.

[0340] Output: Generated album and sharing link.

[0341] What it does: The server creates an album based on the modified photos and videos, and sends a sharing link to the email address of the designated recipient (e.g., grandparents), or sends a notification to a dedicated application.

[0342] Step 6:

[0343] The server generates advertising targets based on the stored data, user profiles, and emotional data, and displays personalized advertisements.

[0344] Input: Modified photos and videos, user emotion data, and user profile data.

[0345] Output: Personalized ads.

[0346] Specific operation: The server generates advertising targets based on the user profile and emotional data, and displays personalized ads using an advertising distribution platform (e.g., Google AdSense).

[0347] This series of processes allows dual-income households to efficiently organize and share photos and videos, and enables companies to effectively deliver personalized advertisements.

[0348] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0349] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0350] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0351] [Second embodiment]

[0352] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0353] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0354] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0355] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0356] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0357] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0358] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0359] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0360] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0361] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0362] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0363] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0364] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken with multiple functions in a server that processes that data.

[0365] 1. Uploading photos and videos

[0366] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0367] 2. Data storage and initial processing

[0368] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0369] 3. AI-powered photo and video selection

[0370] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. This process extracts only high-quality photos and videos from a huge number of photos and videos.

[0371] 4. Automatic correction of background and unwanted reflections

[0372] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0373] 5. Create and share albums

[0374] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0375] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​selects the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0376] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, a user with many travel photos can be served travel-related ads, while a user with many photos of sporting events can be served sporting goods ads. This allows for highly accurate targeting.

[0377] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0378] The processing flow will be explained below.

[0379] Step 1:

[0380] A user takes a photo or video using a smartphone or camera.

[0381] Step 2:

[0382] The user opens a dedicated application and selects the photos and videos they have taken.

[0383] Step 3:

[0384] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0385] Step 4:

[0386] The server receives the uploaded photos and videos.

[0387] Step 5:

[0388] The server stores the received photos and videos in cloud storage.

[0389] Step 6:

[0390] The server extracts metadata from photos and videos (date and time of shooting, location, device information) and stores it in a database.

[0391] Step 7:

[0392] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0393] Step 8:

[0394] An AI algorithm on the server analyzes the photos and videos in the list based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects the best photos and videos.

[0395] Step 9:

[0396] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0397] Step 10:

[0398] The server then saves the edited photos and videos back to cloud storage.

[0399] Step 11:

[0400] The server organizes the selected and edited photos and videos into an album.

[0401] Step 12:

[0402] The server generates an album and sends a sharing notification to the specified recipient's email address or dedicated application.

[0403] Step 13:

[0404] The recipient will receive a link or notification to view the album.

[0405] Step 14:

[0406] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0407] Step 15:

[0408] The server generates targeted advertising data based on user attributes.

[0409] Step 16:

[0410] The server provides the ad targeting data to the relevant companies.

[0411] Example 1

[0412] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0413] In today's dual-income households, it is difficult to efficiently organize and share records of children's growth with family members. To solve this problem, a system that allows users to easily upload, organize, select, edit, and share photos and videos is needed. It also needs the ability to automatically select only high-quality photos and videos from a large number of images and correct backgrounds and unwanted reflections. Furthermore, it is also important to target advertisements based on the recipient's attributes using the shared records.

[0414] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0415] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for saving selected photos and videos in cloud storage and saving the metadata in a database, a means for analyzing and selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos based on the evaluation criteria, a means for compiling the corrected photos and videos into an album and automatically sharing them with designated recipients, a means for sending a sharing link for the album to the recipients via email, and a means for sending a notification to a dedicated application. This allows dual-income households to efficiently organize their children's growth records and easily share them with their families. It also enables advertising targeting based on recipient attributes.

[0416] "Means for uploading photos and videos" refers to a hardware or software interface that provides the functionality for users to send photos and videos they have taken to a server.

[0417] "Means for acquiring and storing metadata" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from uploaded photos and videos, and stores it in a database.

[0418] "Means for saving to cloud storage" refers to the ability to store uploaded photos and videos on a remote server via the Internet, making them accessible.

[0419] The "means of analyzing and selecting based on evaluation criteria" is a function that evaluates photos and videos taken within a specified period of time using criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise, and selects high-quality photos and videos.

[0420] "Automatic correction" is a function that detects unwanted backgrounds and reflections in selected photos and videos and automatically processes them to improve their appearance.

[0421] "A means of organizing photos and videos into an album and automatically sharing them with designated recipients" is a function that lays out edited photos and videos in an album format and shares the album by sending a link or notification to designated recipients.

[0422] The "means for sending a shared link of an album by email" is a function for sending a link of a created album to a recipient's email address.

[0423] The "means for sending a notification to a dedicated application" is a function for sending a notification about the created album to a recipient as a push notification via a dedicated application.

[0424] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system significantly reduces the user's workload by automating the process of uploading, saving, selecting, editing, and sharing photos and videos.

[0425] Uploading photos and videos

[0426] Users upload photos and videos they have taken to the server using a dedicated application (for example, the GrowMemories app). This application provides users with a simple interface and has the ability to select and upload multiple files at once. When uploading, the device compresses the selected photos and videos and sends them to the server.

[0427] Data storage and metadata extraction

[0428] When the server receives uploaded photos and videos, it stores them in cloud storage (e.g., AWS S3). Once storage is complete, the server extracts metadata (such as the date and time of the photo, location, and device information). This metadata is obtained using a dedicated metadata extraction library (e.g., ExifTool). The extracted metadata is stored in a database (e.g., Amazon RDS). If the file format is invalid, the server generates an error message to notify the user.

[0429] AI-powered photo and video selection

[0430] The AI ​​algorithms on the server (e.g., Google Cloud Vision API) analyze photos and videos from a specified period (e.g., the past month) based on criteria such as facial expression, composition, lighting, and the presence or absence of blur and noise, allowing the system to select the highest quality photos and videos.

[0431] Auto-correct photos and videos

[0432] The server uses generative AI (for example, Adobe Photoshop API) to automatically correct backgrounds and unwanted reflections in the selected photos and videos. Automatic correction removes other people and unwanted objects in the background, resulting in better-looking photos and videos.

[0433] Create and share albums

[0434] The server organizes the edited photos and videos into albums, which are automatically generated in HTML or PDF format. The albums are then shared with designated recipients (e.g., grandparents) via a link sent to their email addresses (e.g., using SendGrid) or a notification sent to a dedicated application (e.g., using Firebase Cloud Messaging).

[0435] Specific examples

[0436] For example, consider the case where Mr. A and his family, a dual-income household, use this system. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. After taking the photos, he uploads the data through a dedicated application (the GrowMemories app). The server saves the data in cloud storage, extracts metadata, and stores it in a database. An AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server uses the Adobe Photoshop API to correct unnecessary parts of the background. The server then creates an album based on the corrected photos and videos and sends a link to Mr. A's parents (the child's grandparents) by email. As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow.

[0437] An example of a prompt sentence when using a generative AI model can be written as follows:

[0438] Example prompt:

[0439] "Please explain the programming process of a system that allows dual-income families to efficiently organize and share records of their children's growth. This system has a mechanism whereby users upload photos and videos they have taken, and the data is stored and processed on a server. Please explain in detail how the data will be organized and shared, including the names of the specific hardware and software."

[0440] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0441] Step 1: Upload photos and videos

[0442] The user selects the photos and videos they have taken using a dedicated application. Multiple files can be selected at the same time using the bulk selection function. When the user presses the upload button, the device compresses the selected files and sends them to the server. The input here is the photo or video file they have taken, and the output is the file sent to the server.

[0443] Step 2: Storing data and extracting metadata

[0444] When the server receives uploaded files, it first saves them to cloud storage. It receives uploaded photo and video files as input and saves them to cloud storage as output. After saving is complete, the server extracts metadata for each file. It uses a metadata extraction library to obtain the shooting date and time, location, device information, etc. The input is the saved photo or video file, and the output is the extracted metadata.

[0445] Step 3: Storing Metadata and Error Notification

[0446] The server saves the extracted metadata in a database. If the file format is inappropriate, it sends an error message to the user based on that information. The input is the extracted metadata and the file format check results, and the output is saving the metadata in the database and sending an error message.

[0447] Step 4: AI-powered photo and video selection

[0448] An AI algorithm on the server analyzes photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur and noise. The input is photo and video files from the specified period and their metadata, and the output is a selection of high-quality photos and videos.

[0449] Step 5: Run Auto Fix

[0450] The server uses generative AI to automatically correct backgrounds and unwanted reflections from selected photos and videos. Corrections include removing other people and unnecessary objects. The input is the selected photo or video file, and the output is the corrected photo or video.

[0451] Step 6: Create an album and share it

[0452] The server uses the corrected photos and videos to compile them into an album and generate it in HTML or PDF format. The input is the corrected photos and videos, and the output is the generated album. To share this album with specified recipients, a link is sent to the recipient's email address or a notification is sent to a dedicated application. An email sending library is used to send emails, and a push notification service is used for notifications. The input is the generated album and recipient information, and the output is the shared album.

[0453] Specific examples of operation

[0454] For example, consider the case where Mr. A and his family, a dual-income household, use this system. They go to the zoo on their day off and take 20 photos and 10 videos. After taking the photos, they upload the data to the server using a dedicated application. The server stores the data in cloud storage, extracts metadata, and stores it in a database. Next, an AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server then uses generative AI to correct unwanted background parts. The server then compiles the corrected photos and videos into an album and emails a link to Mr. A's parents. As a result, Mr. A can easily share organized records, and the grandparents can enjoy watching their grandchildren grow up.

[0455] (Application example 1)

[0456] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0457] Today's busy dual-income households make it difficult to organize photos and videos and regularly share them with grandparents and other family members. It's also a tedious task to select high-quality photos and videos from the many available and remove unwanted backgrounds and unwanted reflections. Furthermore, traditional methods for delivering ads tailored to users' interests have limited accuracy, making effective targeting difficult.

[0458] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0459] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients, a means for analyzing specified media data and estimating user interests and concerns, and a means for targeting and delivering advertisements based on the estimation results. This allows dual-income families to efficiently organize and share records of their children's growth, and enables companies to effectively target advertisements.

[0460] "Means for uploading photos and videos" refers to a function that provides an interface for users to transfer photos and videos taken with their smartphones or cameras to a server via the Internet.

[0461] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts metadata such as the date and time of shooting, location, and device information from uploaded media data and stores it in a database.

[0462] "A means of selecting photos and videos from a specified period based on evaluation criteria" is a function that analyzes and selects photos and videos uploaded within a specific period based on criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[0463] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses an AI algorithm to automatically remove or correct unwanted backgrounds and reflections in selected photos and videos.

[0464] "A means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that organizes edited photos and videos into albums and automatically shares them with designated recipients via links or notifications.

[0465] "Means of analyzing specified media data and inferring user interests and concerns" refers to a function that uses an AI algorithm to analyze the content of uploaded photos and videos and predict user interests and concerns.

[0466] "Means for targeting and delivering advertisements based on inferred results" refers to a function that delivers advertisements optimized for individual users based on user profile information obtained from content analysis of photos and videos.

[0467] The system for implementing this invention allows users to upload photos and videos they have taken and process the data on a server to realize advertising targeting based on their interests. The specific configuration and operating procedures of the system are described below.

[0468] System Configuration

[0469] The system mainly includes the following elements:

[0470] 1. User devices: smartphones and digital cameras

[0471] 2. Servers: Cloud storage, AI algorithms, database servers

[0472] 3. Communication Infrastructure: Internet

[0473] Uploading photos and videos

[0474] Users upload photos and videos taken with their smartphones or digital cameras to a server via a dedicated application. The dedicated application provides an interface that allows users to select multiple files at once and easily upload them. The cloud storage used here could be Google Cloud Storage or Amazon S3.

[0475] Retrieving and storing metadata

[0476] The server extracts metadata such as the date and time of the photo or video, location, and device information from the uploaded photo or video, and stores it in a database, typically using PostgreSQL or MySQL.

[0477] Selection of photos and videos within a specified period

[0478] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur or noise.

[0479] Automatic correction of unwanted backgrounds and reflections

[0480] For selected photos and videos, a generative AI model on the server automatically corrects unnecessary backgrounds and reflected elements.

[0481] Create and share albums

[0482] The edited photos and videos are organized into an album, and the server automatically shares them with designated recipients in the form of a link or notification.

[0483] User interest estimation and ad targeting

[0484] The server analyzes the uploaded media data using an AI algorithm to estimate the user's interests and preferences, and then delivers targeted advertisements optimized for the user based on these estimates.

[0485] Specific examples

[0486] For example, a user might upload a photo of a beach they took on vacation. The photo is stored on a server, and AI analyzes the beach image and classifies it as "travel." The server then generates a travel-related ad (e.g., "Special discounts to exotic destinations!") and displays it in the user's app.

[0487] Prompt Sentence Examples

[0488] Here's an example of a prompt to feed to a generative AI model:

[0489] Based on the results of photo analysis, predict user interests and generate appropriate ads.

[0490] In this way, dual-income families can efficiently organize and share records of their children's growth, and companies can perform highly accurate advertising targeting.

[0491] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0492] Step 1:

[0493] Taking and uploading photos and videos

[0494] A user takes photos or videos with a smartphone or digital camera. The captured media data is uploaded to a server via a dedicated application. The input is the captured photos or videos, and the output is the media data stored in cloud storage.

[0495] Step 2:

[0496] Retrieving and storing metadata

[0497] The server extracts metadata (such as the date and time of the photo, location, and device information) from the uploaded photos and videos and stores it in a database. The input is the uploaded photos and videos, and the output is the metadata stored in the database. Specifically, the system uses a software module to extract the metadata.

[0498] Step 3:

[0499] Photo and video selection

[0500] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise). The input is the saved photos and videos, and the output is a selected set of high-quality photos and videos.

[0501] Step 4:

[0502] Correcting unwanted backgrounds and reflections

[0503] The generative AI model in the server automatically corrects unwanted backgrounds and reflected elements in the selected photos and videos. The input is the selected photos and videos, and the output is the corrected photos and videos. The specific operation is to apply a background removal algorithm.

[0504] Step 5:

[0505] Create and share albums

[0506] The server automatically shares the edited photos and videos with the designated recipients by organizing them into albums. The input is the edited photos and videos, and the output is the album link and notification that the recipients can access. Specifically, the album generation software is used.

[0507] Step 6:

[0508] User interest estimation

[0509] The server analyzes the content of the uploaded media data and infers the user's interests. The input is the content information of the uploaded photos and videos, and the output is the inferred results regarding the user's interests. The specific operation is to apply an AI analysis model.

[0510] Step 7:

[0511] Ad Targeting and Delivery

[0512] Advertisements are targeted and delivered based on the inference results. The input is the inference results regarding the user's interests, and the output is targeted advertisements. Specifically, the ad generation module is used to select and deliver appropriate advertisements.

[0513] In this way, by taking specific actions at each step, dual-income families can efficiently organize and share their children's growth records, and companies can achieve highly accurate advertising targeting.

[0514] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0515] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine.

[0516] 1. Uploading photos and videos

[0517] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0518] 2. Data storage and initial processing

[0519] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0520] 3. Photo and video selection using AI and emotion engine

[0521] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. In addition, an emotion engine recognizes the emotions of the subjects in the photos and videos based on facial recognition, and makes selections based on that emotional information. This process extracts high-quality photos and videos that express rich emotions.

[0522] 4. Automatic correction of background and unwanted reflections

[0523] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0524] 5. Create and share albums

[0525] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0526] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0527] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0528] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0529] The processing flow will be explained below.

[0530] Step 1:

[0531] A user takes a photo or video using a smartphone or camera.

[0532] Step 2:

[0533] The user opens the dedicated application and selects the photos and videos they have taken.

[0534] Step 3:

[0535] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0536] Step 4:

[0537] The server receives the uploaded photos and videos and stores them in cloud storage.

[0538] Step 5:

[0539] The server extracts metadata (such as the date and time of the photo, location, and device information) from the photos and videos it receives and stores it in a database.

[0540] Step 6:

[0541] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0542] Step 7:

[0543] An AI algorithm on the server analyzes photos and videos based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects "good" photos and videos.

[0544] Step 8:

[0545] The emotion engine in the server recognizes the faces of the subjects in the selected photos and videos and recognizes their emotions (joy, sadness, surprise, etc.).

[0546] Step 9:

[0547] Based on the emotional information recognized by the emotion engine, the server prioritizes the selection of high-quality photos and videos that express emotions richly.

[0548] Step 10:

[0549] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0550] Step 11:

[0551] The server then saves the edited photos and videos back to cloud storage.

[0552] Step 12:

[0553] The server organizes the edited photos and videos into an album.

[0554] Step 13:

[0555] The server automatically shares the generated album with the designated recipients (e.g. grandparents) by sending a link to their email address or by sending a notification to their application.

[0556] Step 14:

[0557] The recipient will receive a link or notification to view the album.

[0558] Step 15:

[0559] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0560] Step 16:

[0561] The server generates targeted advertising data based on user attributes, and also reflects emotional information recognized by the emotion engine in the targeting data.

[0562] Step 17:

[0563] The server provides the ad targeting data to the relevant companies.

[0564] Example 2

[0565] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0566] It has been difficult for dual-income families to efficiently organize records of their children's growth and regularly share them with grandparents and other family members. Also, sorting and selecting photos and videos taken, and correcting unnecessary backgrounds and reflections, was time-consuming and laborious. Furthermore, it was difficult to select emotionally rich photos and videos.

[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0568] In this invention, the server includes a means for uploading captured photos and videos, a means for storing the uploaded photos and videos in cloud storage and acquiring and storing metadata in a database, a means for selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and an emotion engine, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos using generative AI, and a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients via email or notification. This makes it possible to organize a child's growth record in a high-quality format and easily share it with family members.

[0569] "Means for uploading photos and videos taken" is a function for sending data of photos and videos taken by the user to a server.

[0570] "Means of storing uploaded photos and videos in cloud storage, obtaining metadata, and storing it in a database" refers to a function that stores photos and videos received by the server on the cloud, extracts information (metadata) such as the date and time the file was taken and the location, and stores it in a database.

[0571] "A means of selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and emotion engine" is a function that uses AI to analyze and evaluate photos and videos within a specified period based on criteria such as facial expressions and composition, and then analyzes their emotional state using an emotion engine to select them.

[0572] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos using generative AI" refers to a function that uses AI technology to detect unwanted parts of the background or other objects contained in selected photos and videos, and automatically corrects or removes them.

[0573] "A means of organizing edited photos and videos into an album and automatically sharing them with designated recipients via email or notification" refers to a function that organizes edited data into an album and sends a link to the recipient's email address or automatically shares it via notification.

[0574] This invention is an information processing system that allows dual-income families to efficiently organize their children's growth records and regularly share them with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine and generative AI.

[0575] First, users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once. During uploading, data is transferred securely using SSL / TLS.

[0576] The server then stores the uploaded photos and videos in cloud storage such as Amazon S3. At the same time, it extracts metadata from the photos and videos, such as the date and time of the photo, location, and device information, and stores the metadata in a database (e.g., MySQL). It checks the integrity of the data and notifies the user if the file format is inappropriate.

[0577] Then, an AI algorithm (e.g., a TensorFlow model) stored on the server evaluates photos and videos from a specified period (e.g., the past month). Evaluation criteria include facial expression, composition, lighting, and the presence or absence of blur and noise. An emotion engine is also used to select photos and videos rich in emotion. This results in high-quality photos and videos that express rich expressions and emotions.

[0578] Next, a generative AI (e.g., a GAN model) on the server automatically detects and corrects background and unwanted reflections in the selected photos and videos. For example, it removes other people or unnecessary objects in the background to create a more attractive photo or video.

[0579] Finally, the server automatically organizes the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0580] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A goes to the zoo with his family on a day off and takes photos and videos. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0581] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0582] An example of a prompt might be:

[0583] "Create an album of your best family photos from the past month, removing unwanted backgrounds."

[0584] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0585] Step 1:

[0586] Users select photos and videos they have taken using a dedicated application and upload them to the server through a dedicated interface.

[0587] Input: Photos and video files taken by the user

[0588] Output: Photo and video files uploaded to the server

[0589] Specific example of operation: A user selects all photos and videos of the zoo taken with their smartphone in a dedicated app and clicks the upload button.

[0590] Step 2:

[0591] The server stores the uploaded photos and videos in cloud storage (e.g., Amazon S3), extracts metadata from each file (e.g., date and time of capture, location, device information), and stores it in a database (e.g., MySQL).

[0592] Input: Uploaded photo and video files

[0593] Output: Photos and videos stored in cloud storage, metadata stored in a database

[0594] Specific example of operation: The server saves the received files in Amazon S3 and stores the date, time, and location of the photo in a database.

[0595] Step 3:

[0596] An AI algorithm (e.g., TensorFlow model) on the server analyzes photos and videos taken within a specified period based on evaluation criteria (e.g., facial expression, composition, lighting, presence or absence of blur and noise). The emotion engine analyzes emotions using facial recognition and selects high-quality photos and videos based on this information.

[0597] Input: Photo and video files retrieved from cloud storage, metadata from a database

[0598] Output: High-quality photo and video files selected based on evaluation criteria

[0599] Specific example of operation: The AI ​​model analyzes photos from the past month and selects the family photo with the most smiling faces.

[0600] Step 4:

[0601] Generative AI (e.g., GAN model) within the server automatically detects and corrects unwanted backgrounds and reflections in selected photos and videos.

[0602] Input: Selected photo and video files

[0603] Output: Photo or video file with unwanted background and reflections removed

[0604] Specific example of operation: Removes other people in the background and processes the photo into a clean one.

[0605] Step 5:

[0606] The server organizes the modified photos and videos into an album, and generates links and notifications to automatically share the album with the specified recipients.

[0607] Input: Modified photo or video files

[0608] Output: Digital content in album format, sharing links and notifications

[0609] How it works: The server creates a digital album with the modified family photos and emails the link to the grandparents.

[0610] (Application example 2)

[0611] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0612] In dual-income households, efficiently organizing and sharing the large number of photos and videos taken in a high-quality format is a significant burden. There is also a need to reduce the workload involved in sharing children's growth records with family members, and for companies to effectively target advertisements based on user data. Given this background, a system that easily organizes, edits, and automatically shares photos and videos is needed. Furthermore, a means is needed to efficiently deliver personalized advertisements using stored data and emotional information.

[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0614] In this invention, the server includes means for uploading captured photos and videos, means for acquiring and storing metadata for the uploaded photos and videos, means for selecting photos and videos within a specified period based on evaluation criteria, means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, means for organizing the corrected photos and videos into an album format and automatically sharing them with specified recipients, and means for generating advertising targets based on the stored data and user profiles and displaying personalized advertisements using emotional data. This allows dual-income households to efficiently organize and share photos and videos, and enables companies to easily deliver targeted advertisements based on users' interests.

[0615] - "Means for uploading photos and videos taken" refers to a function that allows users to send photos and videos they have taken to a server via the Internet.

[0616] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from photos and videos received by the server and stores it in a database.

[0617] "A means of selecting photos and videos taken within a specified period based on evaluation criteria" is a function that analyzes photos and videos taken within a specific period based on evaluation criteria such as image quality, composition, facial expression, and lighting conditions, and selects the most suitable ones.

[0618] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses AI to automatically correct unnecessary elements contained in selected photos and videos, improving their quality.

[0619] "Means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that automatically compiles edited photos and videos into albums and electronically shares them with specified recipients.

[0620] "Means for generating advertising targets based on stored data and user profiles, and displaying personalized advertisements using emotional data" refers to a function that generates optimal advertising targets based on the user's stored data, profile information, and emotional monitoring data, and displays them to the user.

[0621] To implement this invention, we need to build a system that efficiently manages and shares photos and videos taken by users. This system consists of major components including smartphones, cloud storage, AI algorithms, emotion engines, and ad distribution platforms.

[0622] System program and processing flow

[0623] 1. Data upload

[0624] Users use a smartphone to upload photos and videos taken using a dedicated application to the server. The uploaded data is stored in cloud storage. During this process, users are provided with a simple interface and can upload multiple files at once.

[0625] 2. Metadata Acquisition and Storage

[0626] Metadata (e.g., shooting date and time, location, device information) is extracted from photos and videos stored in cloud storage and stored in a database. Image processing software (e.g., Google Cloud Vision) is used to obtain the metadata.

[0627] 3. Data selection using AI

[0628] An AI algorithm (e.g., OpenAI GPT-4) analyzes photos and videos from a specified period based on criteria, such as image quality, composition, facial expression, lighting, and the presence or absence of blur or noise, to select high-quality data.

[0629] 4. Emotion Recognition and Data Modification

[0630] An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in photos and videos to obtain emotional data. This emotional data is then reflected in the selection process. Furthermore, generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0631] 5. Album creation and sharing

[0632] The server automatically compiles the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0633] 6. Ad Targeting and Delivery

[0634] Generate advertising targets based on stored data, user profiles, and emotional data. Display personalized ads to the generated advertising targets. This is done using an advertising distribution platform (e.g., Google AdSense).

[0635] Examples of concrete examples and prompts

[0636] As a concrete example, a dual-income household user uploads photos and videos taken at a park on a day off through a dedicated application. The server stores these data in cloud storage, retrieves metadata, and organizes them. AI and an emotion engine select the best photos of the month and correct unwanted background parts. The system then creates an album using the corrected photos and videos, and automatically shares the link with grandparents. Based on the user's profile and emotion data, personalized travel-related advertisements are displayed.

[0637] Examples of prompts:

[0638] Analyze the metadata of photos and videos uploaded by user A over the past month, and use the emotional data obtained from the emotion engine to generate the most appropriate advertising targets.

[0639] Based on the user's emotional data and activity history, create and deliver ads that are likely to be of interest to them. For example, if there are a lot of travel photos and videos, generate travel-related ads, or if there are a lot of photos of sporting events, generate ads for sports equipment.

[0640] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0641] Step 1:

[0642] Users use their smartphones to upload the photos and videos they have taken to the server via a dedicated application.

[0643] Input: Photo and video files taken with a smartphone.

[0644] Output: Photo and video files saved in cloud storage.

[0645] What it does: A user opens the app, selects a photo or video to upload, and the app makes a request to send the file to the server, which receives it and stores it in cloud storage (e.g., AWS S3, Google Cloud Storage).

[0646] Step 2:

[0647] The server extracts metadata (e.g., shooting date and time, location, device information) from photos and videos stored in cloud storage and stores it in a database.

[0648] Input: Photo and video files stored in cloud storage.

[0649] Output: Metadata stored in a database.

[0650] What it does: The server uses image processing software (e.g., Google Cloud Vision) to analyze and extract metadata from photos and videos and store it in a database.

[0651] Step 3:

[0652] The server analyzes photos and videos from a specified period based on evaluation criteria and selects high-quality data.

[0653] Input: Metadata stored in the database and photo and video files.

[0654] Output: A list of selected high-quality photos and videos.

[0655] How it works: The server uses an AI algorithm (e.g., OpenAI GPT-4) to analyze photos and videos from a specified period (e.g., the past month) based on evaluation criteria (image quality, composition, facial expression, lighting conditions, presence or absence of blur and noise), and selects the most suitable ones.

[0656] Step 4:

[0657] The server uses an emotion engine to obtain emotion data from selected photos and videos and corrects the data.

[0658] Input: Selected photos and videos, and user profile information.

[0659] Output: The retouched photo or video files.

[0660] How it works: An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in selected photos and videos to obtain emotional data. Then, a generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0661] Step 5:

[0662] The server organizes the edited photos and videos into an album and automatically shares them with the designated recipients.

[0663] Input: The modified photo or video file.

[0664] Output: Generated album and sharing link.

[0665] What it does: The server creates an album based on the modified photos and videos, and sends a sharing link to the email address of the designated recipient (e.g., grandparents), or sends a notification to a dedicated application.

[0666] Step 6:

[0667] The server generates advertising targets based on the stored data, user profiles, and emotional data, and displays personalized advertisements.

[0668] Input: Modified photos and videos, user emotion data, and user profile data.

[0669] Output: Personalized ads.

[0670] Specific operation: The server generates advertising targets based on the user profile and emotional data, and displays personalized ads using an advertising distribution platform (e.g., Google AdSense).

[0671] This series of processes allows dual-income households to efficiently organize and share photos and videos, and enables companies to effectively deliver personalized advertisements.

[0672] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0673] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0674] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0675] [Third embodiment]

[0676] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0677] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0678] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0679] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0680] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0681] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0682] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0683] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0684] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0685] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0686] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0687] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0688] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken with multiple functions in a server that processes that data.

[0689] 1. Uploading photos and videos

[0690] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0691] 2. Data storage and initial processing

[0692] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0693] 3. AI-powered photo and video selection

[0694] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. This process extracts only high-quality photos and videos from a huge number of photos and videos.

[0695] 4. Automatic correction of background and unwanted reflections

[0696] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0697] 5. Create and share albums

[0698] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0699] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​selects the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0700] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, a user with many travel photos can be served travel-related ads, while a user with many photos of sporting events can be served sporting goods ads. This allows for highly accurate targeting.

[0701] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0702] The processing flow will be explained below.

[0703] Step 1:

[0704] A user takes a photo or video using a smartphone or camera.

[0705] Step 2:

[0706] The user opens a dedicated application and selects the photos and videos they have taken.

[0707] Step 3:

[0708] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0709] Step 4:

[0710] The server receives the uploaded photos and videos.

[0711] Step 5:

[0712] The server stores the received photos and videos in cloud storage.

[0713] Step 6:

[0714] The server extracts metadata from photos and videos (date and time of shooting, location, device information) and stores it in a database.

[0715] Step 7:

[0716] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0717] Step 8:

[0718] An AI algorithm on the server analyzes the photos and videos in the list based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects the best photos and videos.

[0719] Step 9:

[0720] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0721] Step 10:

[0722] The server then saves the edited photos and videos back to cloud storage.

[0723] Step 11:

[0724] The server organizes the selected and edited photos and videos into an album.

[0725] Step 12:

[0726] The server generates an album and sends a sharing notification to the specified recipient's email address or dedicated application.

[0727] Step 13:

[0728] The recipient will receive a link or notification to view the album.

[0729] Step 14:

[0730] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0731] Step 15:

[0732] The server generates targeted advertising data based on user attributes.

[0733] Step 16:

[0734] The server provides the ad targeting data to the relevant companies.

[0735] Example 1

[0736] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0737] In today's dual-income households, it is difficult to efficiently organize and share records of children's growth with family members. To solve this problem, a system that allows users to easily upload, organize, select, edit, and share photos and videos is needed. It also needs the ability to automatically select only high-quality photos and videos from a large number of images and correct backgrounds and unwanted reflections. Furthermore, it is also important to target advertisements based on the recipient's attributes using the shared records.

[0738] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0739] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for saving selected photos and videos in cloud storage and saving the metadata in a database, a means for analyzing and selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos based on the evaluation criteria, a means for compiling the corrected photos and videos into an album and automatically sharing them with designated recipients, a means for sending a sharing link for the album to the recipients via email, and a means for sending a notification to a dedicated application. This allows dual-income households to efficiently organize their children's growth records and easily share them with their families. It also enables advertising targeting based on recipient attributes.

[0740] "Means for uploading photos and videos" refers to a hardware or software interface that provides the functionality for users to send photos and videos they have taken to a server.

[0741] "Means for acquiring and storing metadata" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from uploaded photos and videos, and stores it in a database.

[0742] "Means for saving to cloud storage" refers to the ability to store uploaded photos and videos on a remote server via the Internet, making them accessible.

[0743] The "means of analyzing and selecting based on evaluation criteria" is a function that evaluates photos and videos taken within a specified period of time using criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise, and selects high-quality photos and videos.

[0744] "Automatic correction" is a function that detects unwanted backgrounds and reflections in selected photos and videos and automatically processes them to improve their appearance.

[0745] "A means of organizing photos and videos into an album and automatically sharing them with designated recipients" is a function that lays out edited photos and videos in an album format and shares the album by sending a link or notification to designated recipients.

[0746] The "means for sending a shared link of an album by email" is a function for sending a link of a created album to a recipient's email address.

[0747] The "means for sending a notification to a dedicated application" is a function for sending a notification about the created album to a recipient as a push notification via a dedicated application.

[0748] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system significantly reduces the user's workload by automating the process of uploading, saving, selecting, editing, and sharing photos and videos.

[0749] Uploading photos and videos

[0750] Users upload photos and videos they have taken to the server using a dedicated application (for example, the GrowMemories app). This application provides users with a simple interface and has the ability to select and upload multiple files at once. When uploading, the device compresses the selected photos and videos and sends them to the server.

[0751] Data storage and metadata extraction

[0752] When the server receives uploaded photos and videos, it stores them in cloud storage (e.g., AWS S3). Once storage is complete, the server extracts metadata (such as the date and time of the photo, location, and device information). This metadata is obtained using a dedicated metadata extraction library (e.g., ExifTool). The extracted metadata is stored in a database (e.g., Amazon RDS). If the file format is invalid, the server generates an error message to notify the user.

[0753] AI-powered photo and video selection

[0754] The AI ​​algorithms on the server (e.g., Google Cloud Vision API) analyze photos and videos from a specified period (e.g., the past month) based on criteria such as facial expression, composition, lighting, and the presence or absence of blur and noise, allowing the system to select the highest quality photos and videos.

[0755] Auto-correct photos and videos

[0756] The server uses generative AI (for example, Adobe Photoshop API) to automatically correct backgrounds and unwanted reflections in the selected photos and videos. Automatic correction removes other people and unwanted objects in the background, resulting in better-looking photos and videos.

[0757] Create and share albums

[0758] The server organizes the edited photos and videos into albums, which are automatically generated in HTML or PDF format. The albums are then shared with designated recipients (e.g., grandparents) via a link sent to their email addresses (e.g., using SendGrid) or a notification sent to a dedicated application (e.g., using Firebase Cloud Messaging).

[0759] Specific examples

[0760] For example, consider the case where Mr. A and his family, a dual-income household, use this system. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. After taking the photos, he uploads the data through a dedicated application (the GrowMemories app). The server saves the data in cloud storage, extracts metadata, and stores it in a database. An AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server uses the Adobe Photoshop API to correct unnecessary parts of the background. The server then creates an album based on the corrected photos and videos and sends a link to Mr. A's parents (the child's grandparents) by email. As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow.

[0761] An example of a prompt sentence when using a generative AI model can be written as follows:

[0762] Example prompt:

[0763] "Please explain the programming process of a system that allows dual-income families to efficiently organize and share records of their children's growth. This system has a mechanism whereby users upload photos and videos they have taken, and the data is stored and processed on a server. Please explain in detail how the data will be organized and shared, including the names of the specific hardware and software."

[0764] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0765] Step 1: Upload photos and videos

[0766] The user selects the photos and videos they have taken using a dedicated application. Multiple files can be selected at the same time using the bulk selection function. When the user presses the upload button, the device compresses the selected files and sends them to the server. The input here is the photo or video file they have taken, and the output is the file sent to the server.

[0767] Step 2: Storing data and extracting metadata

[0768] When the server receives uploaded files, it first saves them to cloud storage. It receives uploaded photo and video files as input and saves them to cloud storage as output. After saving is complete, the server extracts metadata for each file. It uses a metadata extraction library to obtain the shooting date and time, location, device information, etc. The input is the saved photo or video file, and the output is the extracted metadata.

[0769] Step 3: Storing Metadata and Error Notification

[0770] The server saves the extracted metadata in a database. If the file format is inappropriate, it sends an error message to the user based on that information. The input is the extracted metadata and the file format check results, and the output is saving the metadata in the database and sending an error message.

[0771] Step 4: AI-powered photo and video selection

[0772] An AI algorithm on the server analyzes photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur and noise. The input is photo and video files from the specified period and their metadata, and the output is a selection of high-quality photos and videos.

[0773] Step 5: Run Auto Fix

[0774] The server uses generative AI to automatically correct backgrounds and unwanted reflections from selected photos and videos. Corrections include removing other people and unnecessary objects. The input is the selected photo or video file, and the output is the corrected photo or video.

[0775] Step 6: Create an album and share it

[0776] The server uses the corrected photos and videos to compile them into an album and generate it in HTML or PDF format. The input is the corrected photos and videos, and the output is the generated album. To share this album with specified recipients, a link is sent to the recipient's email address or a notification is sent to a dedicated application. An email sending library is used to send emails, and a push notification service is used for notifications. The input is the generated album and recipient information, and the output is the shared album.

[0777] Specific examples of operation

[0778] For example, consider the case where Mr. A and his family, a dual-income household, use this system. They go to the zoo on their day off and take 20 photos and 10 videos. After taking the photos, they upload the data to the server using a dedicated application. The server stores the data in cloud storage, extracts metadata, and stores it in a database. Next, an AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server then uses generative AI to correct unwanted background parts. The server then compiles the corrected photos and videos into an album and emails a link to Mr. A's parents. As a result, Mr. A can easily share organized records, and the grandparents can enjoy watching their grandchildren grow up.

[0779] (Application example 1)

[0780] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0781] Today's busy dual-income households make it difficult to organize photos and videos and regularly share them with grandparents and other family members. It's also a tedious task to select high-quality photos and videos from the many available and remove unwanted backgrounds and unwanted reflections. Furthermore, traditional methods for delivering ads tailored to users' interests have limited accuracy, making effective targeting difficult.

[0782] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0783] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients, a means for analyzing specified media data and estimating user interests and concerns, and a means for targeting and delivering advertisements based on the estimation results. This allows dual-income families to efficiently organize and share records of their children's growth, and enables companies to effectively target advertisements.

[0784] "Means for uploading photos and videos" refers to a function that provides an interface for users to transfer photos and videos taken with their smartphones or cameras to a server via the Internet.

[0785] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts metadata such as the date and time of shooting, location, and device information from uploaded media data and stores it in a database.

[0786] "A means of selecting photos and videos from a specified period based on evaluation criteria" is a function that analyzes and selects photos and videos uploaded within a specific period based on criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[0787] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses an AI algorithm to automatically remove or correct unwanted backgrounds and reflections in selected photos and videos.

[0788] "A means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that organizes edited photos and videos into albums and automatically shares them with designated recipients via links or notifications.

[0789] "Means of analyzing specified media data and inferring user interests and concerns" refers to a function that uses an AI algorithm to analyze the content of uploaded photos and videos and predict user interests and concerns.

[0790] "Means for targeting and delivering advertisements based on inferred results" refers to a function that delivers advertisements optimized for individual users based on user profile information obtained from content analysis of photos and videos.

[0791] The system for implementing this invention allows users to upload photos and videos they have taken and process the data on a server to realize advertising targeting based on their interests. The specific configuration and operating procedures of the system are described below.

[0792] System Configuration

[0793] The system mainly includes the following elements:

[0794] 1. User devices: smartphones and digital cameras

[0795] 2. Servers: Cloud storage, AI algorithms, database servers

[0796] 3. Communication Infrastructure: Internet

[0797] Uploading photos and videos

[0798] Users upload photos and videos taken with their smartphones or digital cameras to a server via a dedicated application. The dedicated application provides an interface that allows users to select multiple files at once and easily upload them. The cloud storage used here could be Google Cloud Storage or Amazon S3.

[0799] Retrieving and storing metadata

[0800] The server extracts metadata such as the date and time of the photo or video, location, and device information from the uploaded photo or video, and stores it in a database, typically using PostgreSQL or MySQL.

[0801] Selection of photos and videos within a specified period

[0802] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur or noise.

[0803] Automatic correction of unwanted backgrounds and reflections

[0804] For selected photos and videos, a generative AI model on the server automatically corrects unnecessary backgrounds and reflected elements.

[0805] Create and share albums

[0806] The edited photos and videos are organized into an album, and the server automatically shares them with designated recipients in the form of a link or notification.

[0807] User interest estimation and ad targeting

[0808] The server analyzes the uploaded media data using an AI algorithm to estimate the user's interests and preferences, and then delivers targeted advertisements optimized for the user based on these estimates.

[0809] Specific examples

[0810] For example, a user might upload a photo of a beach they took on vacation. The photo is stored on a server, and AI analyzes the beach image and classifies it as "travel." The server then generates a travel-related ad (e.g., "Special discounts to exotic destinations!") and displays it in the user's app.

[0811] Prompt Sentence Examples

[0812] Here's an example of a prompt to feed to a generative AI model:

[0813] Based on the results of photo analysis, predict user interests and generate appropriate ads.

[0814] In this way, dual-income families can efficiently organize and share records of their children's growth, and companies can perform highly accurate advertising targeting.

[0815] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0816] Step 1:

[0817] Taking and uploading photos and videos

[0818] A user takes photos or videos with a smartphone or digital camera. The captured media data is uploaded to a server via a dedicated application. The input is the captured photos or videos, and the output is the media data stored in cloud storage.

[0819] Step 2:

[0820] Retrieving and storing metadata

[0821] The server extracts metadata (such as the date and time of the photo, location, and device information) from the uploaded photos and videos and stores it in a database. The input is the uploaded photos and videos, and the output is the metadata stored in the database. Specifically, the system uses a software module to extract the metadata.

[0822] Step 3:

[0823] Photo and video selection

[0824] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise). The input is the saved photos and videos, and the output is a selected set of high-quality photos and videos.

[0825] Step 4:

[0826] Correcting unwanted backgrounds and reflections

[0827] The generative AI model in the server automatically corrects unwanted backgrounds and reflected elements in the selected photos and videos. The input is the selected photos and videos, and the output is the corrected photos and videos. The specific operation is to apply a background removal algorithm.

[0828] Step 5:

[0829] Create and share albums

[0830] The server automatically shares the edited photos and videos with the designated recipients by organizing them into albums. The input is the edited photos and videos, and the output is the album link and notification that the recipients can access. Specifically, the album generation software is used.

[0831] Step 6:

[0832] User interest estimation

[0833] The server analyzes the content of the uploaded media data and infers the user's interests. The input is the content information of the uploaded photos and videos, and the output is the inferred results regarding the user's interests. The specific operation is to apply an AI analysis model.

[0834] Step 7:

[0835] Ad Targeting and Delivery

[0836] Advertisements are targeted and delivered based on the inference results. The input is the inference results regarding the user's interests, and the output is targeted advertisements. Specifically, the ad generation module is used to select and deliver appropriate advertisements.

[0837] In this way, by taking specific actions at each step, dual-income families can efficiently organize and share their children's growth records, and companies can achieve highly accurate advertising targeting.

[0838] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0839] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine.

[0840] 1. Uploading photos and videos

[0841] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[0842] 2. Data storage and initial processing

[0843] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[0844] 3. Photo and video selection using AI and emotion engine

[0845] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. In addition, an emotion engine recognizes the emotions of the subjects in the photos and videos based on facial recognition, and makes selections based on that emotional information. This process extracts high-quality photos and videos that express rich emotions.

[0846] 4. Automatic correction of background and unwanted reflections

[0847] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[0848] 5. Create and share albums

[0849] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[0850] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0851] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0852] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[0853] The processing flow will be explained below.

[0854] Step 1:

[0855] A user takes a photo or video using a smartphone or camera.

[0856] Step 2:

[0857] The user opens the dedicated application and selects the photos and videos they have taken.

[0858] Step 3:

[0859] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[0860] Step 4:

[0861] The server receives the uploaded photos and videos and stores them in cloud storage.

[0862] Step 5:

[0863] The server extracts metadata (such as the date and time of the photo, location, and device information) from the photos and videos it receives and stores it in a database.

[0864] Step 6:

[0865] The server generates a list of photos and videos for a specified period (e.g., the past month).

[0866] Step 7:

[0867] An AI algorithm on the server analyzes photos and videos based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects "good" photos and videos.

[0868] Step 8:

[0869] The emotion engine in the server recognizes the faces of the subjects in the selected photos and videos and recognizes their emotions (joy, sadness, surprise, etc.).

[0870] Step 9:

[0871] Based on the emotional information recognized by the emotion engine, the server prioritizes the selection of high-quality photos and videos that express emotions richly.

[0872] Step 10:

[0873] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[0874] Step 11:

[0875] The server then saves the edited photos and videos back to cloud storage.

[0876] Step 12:

[0877] The server organizes the edited photos and videos into an album.

[0878] Step 13:

[0879] The server automatically shares the generated album with the designated recipients (e.g. grandparents) by sending a link to their email address or by sending a notification to their application.

[0880] Step 14:

[0881] The recipient will receive a link or notification to view the album.

[0882] Step 15:

[0883] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[0884] Step 16:

[0885] The server generates targeted advertising data based on user attributes, and also reflects emotional information recognized by the emotion engine in the targeting data.

[0886] Step 17:

[0887] The server provides the ad targeting data to the relevant companies.

[0888] Example 2

[0889] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0890] It has been difficult for dual-income families to efficiently organize records of their children's growth and regularly share them with grandparents and other family members. Also, sorting and selecting photos and videos taken, and correcting unnecessary backgrounds and reflections, was time-consuming and laborious. Furthermore, it was difficult to select emotionally rich photos and videos.

[0891] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0892] In this invention, the server includes a means for uploading captured photos and videos, a means for storing the uploaded photos and videos in cloud storage and acquiring and storing metadata in a database, a means for selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and an emotion engine, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos using generative AI, and a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients via email or notification. This makes it possible to organize a child's growth record in a high-quality format and easily share it with family members.

[0893] "Means for uploading photos and videos taken" is a function for sending data of photos and videos taken by the user to a server.

[0894] "Means of storing uploaded photos and videos in cloud storage, obtaining metadata, and storing it in a database" refers to a function that stores photos and videos received by the server on the cloud, extracts information (metadata) such as the date and time the file was taken and the location, and stores it in a database.

[0895] "A means of selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and emotion engine" is a function that uses AI to analyze and evaluate photos and videos within a specified period based on criteria such as facial expressions and composition, and then analyzes their emotional state using an emotion engine to select them.

[0896] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos using generative AI" refers to a function that uses AI technology to detect unwanted parts of the background or other objects contained in selected photos and videos, and automatically corrects or removes them.

[0897] "A means of organizing edited photos and videos into an album and automatically sharing them with designated recipients via email or notification" refers to a function that organizes edited data into an album and sends a link to the recipient's email address or automatically shares it via notification.

[0898] This invention is an information processing system that allows dual-income families to efficiently organize their children's growth records and regularly share them with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine and generative AI.

[0899] First, users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once. During uploading, data is transferred securely using SSL / TLS.

[0900] The server then stores the uploaded photos and videos in cloud storage such as Amazon S3. At the same time, it extracts metadata from the photos and videos, such as the date and time of the photo, location, and device information, and stores the metadata in a database (e.g., MySQL). It checks the integrity of the data and notifies the user if the file format is inappropriate.

[0901] Then, an AI algorithm (e.g., a TensorFlow model) stored on the server evaluates photos and videos from a specified period (e.g., the past month). Evaluation criteria include facial expression, composition, lighting, and the presence or absence of blur and noise. An emotion engine is also used to select photos and videos rich in emotion. This results in high-quality photos and videos that express rich expressions and emotions.

[0902] Next, a generative AI (e.g., a GAN model) on the server automatically detects and corrects background and unwanted reflections in the selected photos and videos. For example, it removes other people or unnecessary objects in the background to create a more attractive photo or video.

[0903] Finally, the server automatically organizes the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0904] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A goes to the zoo with his family on a day off and takes photos and videos. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[0905] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[0906] An example of a prompt might be:

[0907] "Create an album of your best family photos from the past month, removing unwanted backgrounds."

[0908] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0909] Step 1:

[0910] Users select photos and videos they have taken using a dedicated application and upload them to the server through a dedicated interface.

[0911] Input: Photos and video files taken by the user

[0912] Output: Photo and video files uploaded to the server

[0913] Specific example of operation: A user selects all photos and videos of the zoo taken with their smartphone in a dedicated app and clicks the upload button.

[0914] Step 2:

[0915] The server stores the uploaded photos and videos in cloud storage (e.g., Amazon S3), extracts metadata from each file (e.g., date and time of capture, location, device information), and stores it in a database (e.g., MySQL).

[0916] Input: Uploaded photo and video files

[0917] Output: Photos and videos stored in cloud storage, metadata stored in a database

[0918] Specific example of operation: The server saves the received files in Amazon S3 and stores the date, time, and location of the photo in a database.

[0919] Step 3:

[0920] An AI algorithm (e.g., TensorFlow model) on the server analyzes photos and videos taken within a specified period based on evaluation criteria (e.g., facial expression, composition, lighting, presence or absence of blur and noise). The emotion engine analyzes emotions using facial recognition and selects high-quality photos and videos based on this information.

[0921] Input: Photo and video files retrieved from cloud storage, metadata from a database

[0922] Output: High-quality photo and video files selected based on evaluation criteria

[0923] Specific example of operation: The AI ​​model analyzes photos from the past month and selects the family photo with the most smiling faces.

[0924] Step 4:

[0925] Generative AI (e.g., GAN model) within the server automatically detects and corrects unwanted backgrounds and reflections in selected photos and videos.

[0926] Input: Selected photo and video files

[0927] Output: Photo or video file with unwanted background and reflections removed

[0928] Specific example of operation: Removes other people in the background and processes the photo into a clean one.

[0929] Step 5:

[0930] The server organizes the modified photos and videos into an album, and generates links and notifications to automatically share the album with the specified recipients.

[0931] Input: Modified photo or video files

[0932] Output: Digital content in album format, sharing links and notifications

[0933] How it works: The server creates a digital album with the modified family photos and emails the link to the grandparents.

[0934] (Application example 2)

[0935] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0936] In dual-income households, efficiently organizing and sharing the large number of photos and videos taken in a high-quality format is a significant burden. There is also a need to reduce the workload involved in sharing children's growth records with family members, and for companies to effectively target advertisements based on user data. Given this background, a system that easily organizes, edits, and automatically shares photos and videos is needed. Furthermore, a means is needed to efficiently deliver personalized advertisements using stored data and emotional information.

[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0938] In this invention, the server includes means for uploading captured photos and videos, means for acquiring and storing metadata for the uploaded photos and videos, means for selecting photos and videos within a specified period based on evaluation criteria, means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, means for organizing the corrected photos and videos into an album format and automatically sharing them with specified recipients, and means for generating advertising targets based on the stored data and user profiles and displaying personalized advertisements using emotional data. This allows dual-income households to efficiently organize and share photos and videos, and enables companies to easily deliver targeted advertisements based on users' interests.

[0939] - "Means for uploading photos and videos taken" refers to a function that allows users to send photos and videos they have taken to a server via the Internet.

[0940] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from photos and videos received by the server and stores it in a database.

[0941] "A means of selecting photos and videos taken within a specified period based on evaluation criteria" is a function that analyzes photos and videos taken within a specific period based on evaluation criteria such as image quality, composition, facial expression, and lighting conditions, and selects the most suitable ones.

[0942] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses AI to automatically correct unnecessary elements contained in selected photos and videos, improving their quality.

[0943] "Means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that automatically compiles edited photos and videos into albums and electronically shares them with specified recipients.

[0944] "Means for generating advertising targets based on stored data and user profiles, and displaying personalized advertisements using emotional data" refers to a function that generates optimal advertising targets based on the user's stored data, profile information, and emotional monitoring data, and displays them to the user.

[0945] To implement this invention, we need to build a system that efficiently manages and shares photos and videos taken by users. This system consists of major components including smartphones, cloud storage, AI algorithms, emotion engines, and ad distribution platforms.

[0946] System program and processing flow

[0947] 1. Data upload

[0948] Users use a smartphone to upload photos and videos taken using a dedicated application to the server. The uploaded data is stored in cloud storage. During this process, users are provided with a simple interface and can upload multiple files at once.

[0949] 2. Metadata Acquisition and Storage

[0950] Metadata (e.g., shooting date and time, location, device information) is extracted from photos and videos stored in cloud storage and stored in a database. Image processing software (e.g., Google Cloud Vision) is used to obtain the metadata.

[0951] 3. Data selection using AI

[0952] An AI algorithm (e.g., OpenAI GPT-4) analyzes photos and videos from a specified period based on criteria, such as image quality, composition, facial expression, lighting, and the presence or absence of blur or noise, to select high-quality data.

[0953] 4. Emotion Recognition and Data Modification

[0954] An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in photos and videos to obtain emotional data. This emotional data is then reflected in the selection process. Furthermore, generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0955] 5. Album creation and sharing

[0956] The server automatically compiles the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[0957] 6. Ad Targeting and Delivery

[0958] Generate advertising targets based on stored data, user profiles, and emotional data. Display personalized ads to the generated advertising targets. This is done using an advertising distribution platform (e.g., Google AdSense).

[0959] Examples of concrete examples and prompts

[0960] As a concrete example, a dual-income household user uploads photos and videos taken at a park on a day off through a dedicated application. The server stores these data in cloud storage, retrieves metadata, and organizes them. AI and an emotion engine select the best photos of the month and correct unwanted background parts. The system then creates an album using the corrected photos and videos, and automatically shares the link with grandparents. Based on the user's profile and emotion data, personalized travel-related advertisements are displayed.

[0961] Examples of prompts:

[0962] Analyze the metadata of photos and videos uploaded by user A over the past month, and use the emotional data obtained from the emotion engine to generate the most appropriate advertising targets.

[0963] Based on the user's emotional data and activity history, create and deliver ads that are likely to be of interest to them. For example, if there are a lot of travel photos and videos, generate travel-related ads, or if there are a lot of photos of sporting events, generate ads for sports equipment.

[0964] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0965] Step 1:

[0966] Users use their smartphones to upload the photos and videos they have taken to the server via a dedicated application.

[0967] Input: Photo and video files taken with a smartphone.

[0968] Output: Photo and video files saved in cloud storage.

[0969] What it does: A user opens the app, selects a photo or video to upload, and the app makes a request to send the file to the server, which receives it and stores it in cloud storage (e.g., AWS S3, Google Cloud Storage).

[0970] Step 2:

[0971] The server extracts metadata (e.g., shooting date and time, location, device information) from photos and videos stored in cloud storage and stores it in a database.

[0972] Input: Photo and video files stored in cloud storage.

[0973] Output: Metadata stored in a database.

[0974] What it does: The server uses image processing software (e.g., Google Cloud Vision) to analyze and extract metadata from photos and videos and store it in a database.

[0975] Step 3:

[0976] The server analyzes photos and videos from a specified period based on evaluation criteria and selects high-quality data.

[0977] Input: Metadata stored in the database and photo and video files.

[0978] Output: A list of selected high-quality photos and videos.

[0979] How it works: The server uses an AI algorithm (e.g., OpenAI GPT-4) to analyze photos and videos from a specified period (e.g., the past month) based on evaluation criteria (image quality, composition, facial expression, lighting conditions, presence or absence of blur and noise), and selects the most suitable ones.

[0980] Step 4:

[0981] The server uses an emotion engine to obtain emotion data from selected photos and videos and corrects the data.

[0982] Input: Selected photos and videos, and user profile information.

[0983] Output: The retouched photo or video files.

[0984] How it works: An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in selected photos and videos to obtain emotional data. Then, a generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[0985] Step 5:

[0986] The server organizes the edited photos and videos into an album and automatically shares them with the designated recipients.

[0987] Input: The modified photo or video file.

[0988] Output: Generated album and sharing link.

[0989] What it does: The server creates an album based on the modified photos and videos, and sends a sharing link to the email address of the designated recipient (e.g., grandparents), or sends a notification to a dedicated application.

[0990] Step 6:

[0991] The server generates advertising targets based on the stored data, user profiles, and emotional data, and displays personalized advertisements.

[0992] Input: Modified photos and videos, user emotion data, and user profile data.

[0993] Output: Personalized ads.

[0994] Specific operation: The server generates advertising targets based on the user profile and emotional data, and displays personalized ads using an advertising distribution platform (e.g., Google AdSense).

[0995] This series of processes allows dual-income households to efficiently organize and share photos and videos, and enables companies to effectively deliver personalized advertisements.

[0996] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0997] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0998] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0999] [Fourth embodiment]

[1000] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1001] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1002] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1003] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1004] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1005] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1006] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1007] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1008] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1009] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1010] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1011] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1012] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1013] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken with multiple functions in a server that processes that data.

[1014] 1. Uploading photos and videos

[1015] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[1016] 2. Data storage and initial processing

[1017] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[1018] 3. AI-powered photo and video selection

[1019] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. This process extracts only high-quality photos and videos from a huge number of photos and videos.

[1020] 4. Automatic correction of background and unwanted reflections

[1021] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[1022] 5. Create and share albums

[1023] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[1024] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​selects the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[1025] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, a user with many travel photos can be served travel-related ads, while a user with many photos of sporting events can be served sporting goods ads. This allows for highly accurate targeting.

[1026] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[1027] The processing flow will be explained below.

[1028] Step 1:

[1029] A user takes a photo or video using a smartphone or camera.

[1030] Step 2:

[1031] The user opens a dedicated application and selects the photos and videos they have taken.

[1032] Step 3:

[1033] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[1034] Step 4:

[1035] The server receives the uploaded photos and videos.

[1036] Step 5:

[1037] The server stores the received photos and videos in cloud storage.

[1038] Step 6:

[1039] The server extracts metadata from photos and videos (date and time of shooting, location, device information) and stores it in a database.

[1040] Step 7:

[1041] The server generates a list of photos and videos for a specified period (e.g., the past month).

[1042] Step 8:

[1043] An AI algorithm on the server analyzes the photos and videos in the list based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects the best photos and videos.

[1044] Step 9:

[1045] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[1046] Step 10:

[1047] The server then saves the edited photos and videos back to cloud storage.

[1048] Step 11:

[1049] The server organizes the selected and edited photos and videos into an album.

[1050] Step 12:

[1051] The server generates an album and sends a sharing notification to the specified recipient's email address or dedicated application.

[1052] Step 13:

[1053] The recipient will receive a link or notification to view the album.

[1054] Step 14:

[1055] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[1056] Step 15:

[1057] The server generates targeted advertising data based on user attributes.

[1058] Step 16:

[1059] The server provides the ad targeting data to the relevant companies.

[1060] Example 1

[1061] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1062] In today's dual-income households, it is difficult to efficiently organize and share records of children's growth with family members. To solve this problem, a system that allows users to easily upload, organize, select, edit, and share photos and videos is needed. It also needs the ability to automatically select only high-quality photos and videos from a large number of images and correct backgrounds and unwanted reflections. Furthermore, it is also important to target advertisements based on the recipient's attributes using the shared records.

[1063] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1064] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for saving selected photos and videos in cloud storage and saving the metadata in a database, a means for analyzing and selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos based on the evaluation criteria, a means for compiling the corrected photos and videos into an album and automatically sharing them with designated recipients, a means for sending a sharing link for the album to the recipients via email, and a means for sending a notification to a dedicated application. This allows dual-income households to efficiently organize their children's growth records and easily share them with their families. It also enables advertising targeting based on recipient attributes.

[1065] "Means for uploading photos and videos" refers to a hardware or software interface that provides the functionality for users to send photos and videos they have taken to a server.

[1066] "Means for acquiring and storing metadata" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from uploaded photos and videos, and stores it in a database.

[1067] "Means for saving to cloud storage" refers to the ability to store uploaded photos and videos on a remote server via the Internet, making them accessible.

[1068] The "means of analyzing and selecting based on evaluation criteria" is a function that evaluates photos and videos taken within a specified period of time using criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise, and selects high-quality photos and videos.

[1069] "Automatic correction" is a function that detects unwanted backgrounds and reflections in selected photos and videos and automatically processes them to improve their appearance.

[1070] "A means of organizing photos and videos into an album and automatically sharing them with designated recipients" is a function that lays out edited photos and videos in an album format and shares the album by sending a link or notification to designated recipients.

[1071] The "means for sending a shared link of an album by email" is a function for sending a link of a created album to a recipient's email address.

[1072] The "means for sending a notification to a dedicated application" is a function for sending a notification about the created album to a recipient as a push notification via a dedicated application.

[1073] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system significantly reduces the user's workload by automating the process of uploading, saving, selecting, editing, and sharing photos and videos.

[1074] Uploading photos and videos

[1075] Users upload photos and videos they have taken to the server using a dedicated application (for example, the GrowMemories app). This application provides users with a simple interface and has the ability to select and upload multiple files at once. When uploading, the device compresses the selected photos and videos and sends them to the server.

[1076] Data storage and metadata extraction

[1077] When the server receives uploaded photos and videos, it stores them in cloud storage (e.g., AWS S3). Once storage is complete, the server extracts metadata (such as the date and time of the photo, location, and device information). This metadata is obtained using a dedicated metadata extraction library (e.g., ExifTool). The extracted metadata is stored in a database (e.g., Amazon RDS). If the file format is invalid, the server generates an error message to notify the user.

[1078] AI-powered photo and video selection

[1079] The AI ​​algorithms on the server (e.g., Google Cloud Vision API) analyze photos and videos from a specified period (e.g., the past month) based on criteria such as facial expression, composition, lighting, and the presence or absence of blur and noise, allowing the system to select the highest quality photos and videos.

[1080] Auto-correct photos and videos

[1081] The server uses generative AI (for example, Adobe Photoshop API) to automatically correct backgrounds and unwanted reflections in the selected photos and videos. Automatic correction removes other people and unwanted objects in the background, resulting in better-looking photos and videos.

[1082] Create and share albums

[1083] The server organizes the edited photos and videos into albums, which are automatically generated in HTML or PDF format. The albums are then shared with designated recipients (e.g., grandparents) via a link sent to their email addresses (e.g., using SendGrid) or a notification sent to a dedicated application (e.g., using Firebase Cloud Messaging).

[1084] Specific examples

[1085] For example, consider the case where Mr. A and his family, a dual-income household, use this system. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. After taking the photos, he uploads the data through a dedicated application (the GrowMemories app). The server saves the data in cloud storage, extracts metadata, and stores it in a database. An AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server uses the Adobe Photoshop API to correct unnecessary parts of the background. The server then creates an album based on the corrected photos and videos and sends a link to Mr. A's parents (the child's grandparents) by email. As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow.

[1086] An example of a prompt sentence when using a generative AI model can be written as follows:

[1087] Example prompt:

[1088] "Please explain the programming process of a system that allows dual-income families to efficiently organize and share records of their children's growth. This system has a mechanism whereby users upload photos and videos they have taken, and the data is stored and processed on a server. Please explain in detail how the data will be organized and shared, including the names of the specific hardware and software."

[1089] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1090] Step 1: Upload photos and videos

[1091] The user selects the photos and videos they have taken using a dedicated application. Multiple files can be selected at the same time using the bulk selection function. When the user presses the upload button, the device compresses the selected files and sends them to the server. The input here is the photo or video file they have taken, and the output is the file sent to the server.

[1092] Step 2: Storing data and extracting metadata

[1093] When the server receives uploaded files, it first saves them to cloud storage. It receives uploaded photo and video files as input and saves them to cloud storage as output. After saving is complete, the server extracts metadata for each file. It uses a metadata extraction library to obtain the shooting date and time, location, device information, etc. The input is the saved photo or video file, and the output is the extracted metadata.

[1094] Step 3: Storing Metadata and Error Notification

[1095] The server saves the extracted metadata in a database. If the file format is inappropriate, it sends an error message to the user based on that information. The input is the extracted metadata and the file format check results, and the output is saving the metadata in the database and sending an error message.

[1096] Step 4: AI-powered photo and video selection

[1097] An AI algorithm on the server analyzes photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur and noise. The input is photo and video files from the specified period and their metadata, and the output is a selection of high-quality photos and videos.

[1098] Step 5: Run Auto Fix

[1099] The server uses generative AI to automatically correct backgrounds and unwanted reflections from selected photos and videos. Corrections include removing other people and unnecessary objects. The input is the selected photo or video file, and the output is the corrected photo or video.

[1100] Step 6: Create an album and share it

[1101] The server uses the corrected photos and videos to compile them into an album and generate it in HTML or PDF format. The input is the corrected photos and videos, and the output is the generated album. To share this album with specified recipients, a link is sent to the recipient's email address or a notification is sent to a dedicated application. An email sending library is used to send emails, and a push notification service is used for notifications. The input is the generated album and recipient information, and the output is the shared album.

[1102] Specific examples of operation

[1103] For example, consider the case where Mr. A and his family, a dual-income household, use this system. They go to the zoo on their day off and take 20 photos and 10 videos. After taking the photos, they upload the data to the server using a dedicated application. The server stores the data in cloud storage, extracts metadata, and stores it in a database. Next, an AI algorithm evaluates the photos and videos and selects the five photos and three videos with the highest ratings. The server then uses generative AI to correct unwanted background parts. The server then compiles the corrected photos and videos into an album and emails a link to Mr. A's parents. As a result, Mr. A can easily share organized records, and the grandparents can enjoy watching their grandchildren grow up.

[1104] (Application example 1)

[1105] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1106] Today's busy dual-income households make it difficult to organize photos and videos and regularly share them with grandparents and other family members. It's also a tedious task to select high-quality photos and videos from the many available and remove unwanted backgrounds and unwanted reflections. Furthermore, traditional methods for delivering ads tailored to users' interests have limited accuracy, making effective targeting difficult.

[1107] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1108] In this invention, the server includes a means for uploading captured photos and videos, a means for acquiring and saving metadata for the uploaded photos and videos, a means for selecting photos and videos from a specified period based on evaluation criteria, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients, a means for analyzing specified media data and estimating user interests and concerns, and a means for targeting and delivering advertisements based on the estimation results. This allows dual-income families to efficiently organize and share records of their children's growth, and enables companies to effectively target advertisements.

[1109] "Means for uploading photos and videos" refers to a function that provides an interface for users to transfer photos and videos taken with their smartphones or cameras to a server via the Internet.

[1110] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts metadata such as the date and time of shooting, location, and device information from uploaded media data and stores it in a database.

[1111] "A means of selecting photos and videos from a specified period based on evaluation criteria" is a function that analyzes and selects photos and videos uploaded within a specific period based on criteria such as facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[1112] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses an AI algorithm to automatically remove or correct unwanted backgrounds and reflections in selected photos and videos.

[1113] "A means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that organizes edited photos and videos into albums and automatically shares them with designated recipients via links or notifications.

[1114] "Means of analyzing specified media data and inferring user interests and concerns" refers to a function that uses an AI algorithm to analyze the content of uploaded photos and videos and predict user interests and concerns.

[1115] "Means for targeting and delivering advertisements based on inferred results" refers to a function that delivers advertisements optimized for individual users based on user profile information obtained from content analysis of photos and videos.

[1116] The system for implementing this invention allows users to upload photos and videos they have taken and process the data on a server to realize advertising targeting based on their interests. The specific configuration and operating procedures of the system are described below.

[1117] System Configuration

[1118] The system mainly includes the following elements:

[1119] 1. User devices: smartphones and digital cameras

[1120] 2. Servers: Cloud storage, AI algorithms, database servers

[1121] 3. Communication Infrastructure: Internet

[1122] Uploading photos and videos

[1123] Users upload photos and videos taken with their smartphones or digital cameras to a server via a dedicated application. The dedicated application provides an interface that allows users to select multiple files at once and easily upload them. The cloud storage used here could be Google Cloud Storage or Amazon S3.

[1124] Retrieving and storing metadata

[1125] The server extracts metadata such as the date and time of the photo or video, location, and device information from the uploaded photo or video, and stores it in a database, typically using PostgreSQL or MySQL.

[1126] Selection of photos and videos within a specified period

[1127] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria, including facial expression, composition, lighting, and the presence or absence of blur or noise.

[1128] Automatic correction of unwanted backgrounds and reflections

[1129] For selected photos and videos, a generative AI model on the server automatically corrects unnecessary backgrounds and reflected elements.

[1130] Create and share albums

[1131] The edited photos and videos are organized into an album, and the server automatically shares them with designated recipients in the form of a link or notification.

[1132] User interest estimation and ad targeting

[1133] The server analyzes the uploaded media data using an AI algorithm to estimate the user's interests and preferences, and then delivers targeted advertisements optimized for the user based on these estimates.

[1134] Specific examples

[1135] For example, a user might upload a photo of a beach they took on vacation. The photo is stored on a server, and AI analyzes the beach image and classifies it as "travel." The server then generates a travel-related ad (e.g., "Special discounts to exotic destinations!") and displays it in the user's app.

[1136] Prompt Sentence Examples

[1137] Here's an example of a prompt to feed to a generative AI model:

[1138] Based on the results of photo analysis, predict user interests and generate appropriate ads.

[1139] In this way, dual-income families can efficiently organize and share records of their children's growth, and companies can perform highly accurate advertising targeting.

[1140] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1141] Step 1:

[1142] Taking and uploading photos and videos

[1143] A user takes photos or videos with a smartphone or digital camera. The captured media data is uploaded to a server via a dedicated application. The input is the captured photos or videos, and the output is the media data stored in cloud storage.

[1144] Step 2:

[1145] Retrieving and storing metadata

[1146] The server extracts metadata (such as the date and time of the photo, location, and device information) from the uploaded photos and videos and stores it in a database. The input is the uploaded photos and videos, and the output is the metadata stored in the database. Specifically, the system uses a software module to extract the metadata.

[1147] Step 3:

[1148] Photo and video selection

[1149] An AI algorithm on the server selects photos and videos from a specified period based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise). The input is the saved photos and videos, and the output is a selected set of high-quality photos and videos.

[1150] Step 4:

[1151] Correcting unwanted backgrounds and reflections

[1152] The generative AI model in the server automatically corrects unwanted backgrounds and reflected elements in the selected photos and videos. The input is the selected photos and videos, and the output is the corrected photos and videos. The specific operation is to apply a background removal algorithm.

[1153] Step 5:

[1154] Create and share albums

[1155] The server automatically shares the edited photos and videos with the designated recipients by organizing them into albums. The input is the edited photos and videos, and the output is the album link and notification that the recipients can access. Specifically, the album generation software is used.

[1156] Step 6:

[1157] User interest estimation

[1158] The server analyzes the content of the uploaded media data and infers the user's interests. The input is the content information of the uploaded photos and videos, and the output is the inferred results regarding the user's interests. The specific operation is to apply an AI analysis model.

[1159] Step 7:

[1160] Ad Targeting and Delivery

[1161] Advertisements are targeted and delivered based on the inference results. The input is the inference results regarding the user's interests, and the output is targeted advertisements. Specifically, the ad generation module is used to select and deliver appropriate advertisements.

[1162] In this way, by taking specific actions at each step, dual-income families can efficiently organize and share their children's growth records, and companies can achieve highly accurate advertising targeting.

[1163] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1164] This invention is a system that allows dual-income families to efficiently organize and regularly share their children's growth records with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine.

[1165] 1. Uploading photos and videos

[1166] Users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once.

[1167] 2. Data storage and initial processing

[1168] The server stores the uploaded photos and videos in cloud storage, extracts the necessary metadata (e.g., date and time of the photo, location, device information) and stores it in a database. At this stage, the server checks the integrity of the data and notifies the user if the file format is inappropriate.

[1169] 3. Photo and video selection using AI and emotion engine

[1170] An AI algorithm on the server analyzes photos and videos from a specified period (e.g., the past month) based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur and noise) and selects "good" photos and videos. In addition, an emotion engine recognizes the emotions of the subjects in the photos and videos based on facial recognition, and makes selections based on that emotional information. This process extracts high-quality photos and videos that express rich emotions.

[1171] 4. Automatic correction of background and unwanted reflections

[1172] The generative AI on the server automatically detects and corrects background and unwanted reflections in selected photos and videos, for example, removing other people or unnecessary objects in the background to create better-looking photos and videos.

[1173] 5. Create and share albums

[1174] The server automatically organizes the edited photos and videos into albums, which are then automatically shared with designated recipients (e.g., grandparents) via a link sent to their email addresses or via a notification sent to a dedicated application.

[1175] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A takes many photos and videos when he and his family go to the zoo on a holiday. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[1176] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[1177] This system not only allows dual-income families to efficiently organize and share records of their children's growth, but also enables companies to effectively target their advertisements.

[1178] The processing flow will be explained below.

[1179] Step 1:

[1180] A user takes a photo or video using a smartphone or camera.

[1181] Step 2:

[1182] The user opens the dedicated application and selects the photos and videos they have taken.

[1183] Step 3:

[1184] When the user presses the "Upload" button, the device sends the selected photos and videos to the server.

[1185] Step 4:

[1186] The server receives the uploaded photos and videos and stores them in cloud storage.

[1187] Step 5:

[1188] The server extracts metadata (such as the date and time of the photo, location, and device information) from the photos and videos it receives and stores it in a database.

[1189] Step 6:

[1190] The server generates a list of photos and videos for a specified period (e.g., the past month).

[1191] Step 7:

[1192] An AI algorithm on the server analyzes photos and videos based on evaluation criteria (facial expression, composition, lighting, presence or absence of blur or noise) and selects "good" photos and videos.

[1193] Step 8:

[1194] The emotion engine in the server recognizes the faces of the subjects in the selected photos and videos and recognizes their emotions (joy, sadness, surprise, etc.).

[1195] Step 9:

[1196] Based on the emotional information recognized by the emotion engine, the server prioritizes the selection of high-quality photos and videos that express emotions richly.

[1197] Step 10:

[1198] The generative AI on the server automatically detects and corrects backgrounds and unwanted reflections in selected photos and videos.

[1199] Step 11:

[1200] The server then saves the edited photos and videos back to cloud storage.

[1201] Step 12:

[1202] The server organizes the edited photos and videos into an album.

[1203] Step 13:

[1204] The server automatically shares the generated album with the designated recipients (e.g. grandparents) by sending a link to their email address or by sending a notification to their application.

[1205] Step 14:

[1206] The recipient will receive a link or notification to view the album.

[1207] Step 15:

[1208] AI on the server analyzes the stored data and classifies user attributes (photo frequency, content, location, etc.).

[1209] Step 16:

[1210] The server generates targeted advertising data based on user attributes, and also reflects emotional information recognized by the emotion engine in the targeting data.

[1211] Step 17:

[1212] The server provides the ad targeting data to the relevant companies.

[1213] Example 2

[1214] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1215] It has been difficult for dual-income families to efficiently organize records of their children's growth and regularly share them with grandparents and other family members. Also, sorting and selecting photos and videos taken, and correcting unnecessary backgrounds and reflections, was time-consuming and laborious. Furthermore, it was difficult to select emotionally rich photos and videos.

[1216] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1217] In this invention, the server includes a means for uploading captured photos and videos, a means for storing the uploaded photos and videos in cloud storage and acquiring and storing metadata in a database, a means for selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and an emotion engine, a means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos using generative AI, and a means for compiling the corrected photos and videos into an album format and automatically sharing them with specified recipients via email or notification. This makes it possible to organize a child's growth record in a high-quality format and easily share it with family members.

[1218] "Means for uploading photos and videos taken" is a function for sending data of photos and videos taken by the user to a server.

[1219] "Means of storing uploaded photos and videos in cloud storage, obtaining metadata, and storing it in a database" refers to a function that stores photos and videos received by the server on the cloud, extracts information (metadata) such as the date and time the file was taken and the location, and stores it in a database.

[1220] "A means of selecting photos and videos within a specified period based on evaluation criteria using an AI algorithm and emotion engine" is a function that uses AI to analyze and evaluate photos and videos within a specified period based on criteria such as facial expressions and composition, and then analyzes their emotional state using an emotion engine to select them.

[1221] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos using generative AI" refers to a function that uses AI technology to detect unwanted parts of the background or other objects contained in selected photos and videos, and automatically corrects or removes them.

[1222] "A means of organizing edited photos and videos into an album and automatically sharing them with designated recipients via email or notification" refers to a function that organizes edited data into an album and sends a link to the recipient's email address or automatically shares it via notification.

[1223] This invention is an information processing system that allows dual-income families to efficiently organize their children's growth records and regularly share them with grandparents and other family members. This system is realized by combining a means for users to upload photos and videos they have taken, multiple functions in the server that process that data, and an emotion engine and generative AI.

[1224] First, users upload photos and videos taken with their smartphones or cameras to a server using a dedicated application. The application provides users with a simple interface that allows them to select and upload multiple files at once. During uploading, data is transferred securely using SSL / TLS.

[1225] The server then stores the uploaded photos and videos in cloud storage such as Amazon S3. At the same time, it extracts metadata from the photos and videos, such as the date and time of the photo, location, and device information, and stores the metadata in a database (e.g., MySQL). It checks the integrity of the data and notifies the user if the file format is inappropriate.

[1226] Then, an AI algorithm (e.g., a TensorFlow model) stored on the server evaluates photos and videos from a specified period (e.g., the past month). Evaluation criteria include facial expression, composition, lighting, and the presence or absence of blur and noise. An emotion engine is also used to select photos and videos rich in emotion. This results in high-quality photos and videos that express rich expressions and emotions.

[1227] Next, a generative AI (e.g., a GAN model) on the server automatically detects and corrects background and unwanted reflections in the selected photos and videos. For example, it removes other people or unnecessary objects in the background to create a more attractive photo or video.

[1228] Finally, the server automatically organizes the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[1229] As a concrete example, consider the case where this system is used by Mr. A's family, a dual-income household. Mr. A goes to the zoo with his family on a day off and takes photos and videos. He then uploads the captured data through a dedicated application. The server stores this data on the cloud, obtains metadata, and organizes it. The AI ​​and emotion engine select the best shots of the month and corrects unnecessary parts of the background. An album is then created using the corrected photos and videos, and the link is automatically shared with Mr. A's parents (the child's grandparents). As a result, Mr. A is freed from the tedious task of organizing, and the grandparents can regularly enjoy watching their grandchildren grow up.

[1230] For B2B applications, advertising targeting can also be performed based on information obtained from stored data and user profiles. For example, travel-related ads can be delivered to a user with many travel photos, while sports goods ads can be displayed to a user with many photos of sporting events. In this case, the emotional data obtained by the emotion engine is also reflected in the targeting data, enabling more accurate advertising delivery.

[1231] An example of a prompt might be:

[1232] "Create an album of your best family photos from the past month, removing unwanted backgrounds."

[1233] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1234] Step 1:

[1235] Users select photos and videos they have taken using a dedicated application and upload them to the server through a dedicated interface.

[1236] Input: Photos and video files taken by the user

[1237] Output: Photo and video files uploaded to the server

[1238] Specific example of operation: A user selects all photos and videos of the zoo taken with their smartphone in a dedicated app and clicks the upload button.

[1239] Step 2:

[1240] The server stores the uploaded photos and videos in cloud storage (e.g., Amazon S3), extracts metadata from each file (e.g., date and time of capture, location, device information), and stores it in a database (e.g., MySQL).

[1241] Input: Uploaded photo and video files

[1242] Output: Photos and videos stored in cloud storage, metadata stored in a database

[1243] Specific example of operation: The server saves the received files in Amazon S3 and stores the date, time, and location of the photo in a database.

[1244] Step 3:

[1245] An AI algorithm (e.g., TensorFlow model) on the server analyzes photos and videos taken within a specified period based on evaluation criteria (e.g., facial expression, composition, lighting, presence or absence of blur and noise). The emotion engine analyzes emotions using facial recognition and selects high-quality photos and videos based on this information.

[1246] Input: Photo and video files retrieved from cloud storage, metadata from a database

[1247] Output: High-quality photo and video files selected based on evaluation criteria

[1248] Specific example of operation: The AI ​​model analyzes photos from the past month and selects the family photo with the most smiling faces.

[1249] Step 4:

[1250] Generative AI (e.g., GAN model) within the server automatically detects and corrects unwanted backgrounds and reflections in selected photos and videos.

[1251] Input: Selected photo and video files

[1252] Output: Photo or video file with unwanted background and reflections removed

[1253] Specific example of operation: Removes other people in the background and processes the photo into a clean one.

[1254] Step 5:

[1255] The server organizes the modified photos and videos into an album, and generates links and notifications to automatically share the album with the specified recipients.

[1256] Input: Modified photo or video files

[1257] Output: Digital content in album format, sharing links and notifications

[1258] How it works: The server creates a digital album with the modified family photos and emails the link to the grandparents.

[1259] (Application example 2)

[1260] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1261] In dual-income households, efficiently organizing and sharing the large number of photos and videos taken in a high-quality format is a significant burden. There is also a need to reduce the workload involved in sharing children's growth records with family members, and for companies to effectively target advertisements based on user data. Given this background, a system that easily organizes, edits, and automatically shares photos and videos is needed. Furthermore, a means is needed to efficiently deliver personalized advertisements using stored data and emotional information.

[1262] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1263] In this invention, the server includes means for uploading captured photos and videos, means for acquiring and storing metadata for the uploaded photos and videos, means for selecting photos and videos within a specified period based on evaluation criteria, means for automatically correcting unwanted backgrounds and reflections in the selected photos and videos, means for organizing the corrected photos and videos into an album format and automatically sharing them with specified recipients, and means for generating advertising targets based on the stored data and user profiles and displaying personalized advertisements using emotional data. This allows dual-income households to efficiently organize and share photos and videos, and enables companies to easily deliver targeted advertisements based on users' interests.

[1264] - "Means for uploading photos and videos taken" refers to a function that allows users to send photos and videos they have taken to a server via the Internet.

[1265] "Means for obtaining and storing metadata of uploaded photos and videos" refers to a function that extracts additional information such as the date and time of shooting, location, and device information from photos and videos received by the server and stores it in a database.

[1266] "A means of selecting photos and videos taken within a specified period based on evaluation criteria" is a function that analyzes photos and videos taken within a specific period based on evaluation criteria such as image quality, composition, facial expression, and lighting conditions, and selects the most suitable ones.

[1267] "Means for automatically correcting unwanted backgrounds and reflections in selected photos and videos" is a function that uses AI to automatically correct unnecessary elements contained in selected photos and videos, improving their quality.

[1268] "Means of organizing edited photos and videos into albums and automatically sharing them with designated recipients" refers to a function that automatically compiles edited photos and videos into albums and electronically shares them with specified recipients.

[1269] "Means for generating advertising targets based on stored data and user profiles, and displaying personalized advertisements using emotional data" refers to a function that generates optimal advertising targets based on the user's stored data, profile information, and emotional monitoring data, and displays them to the user.

[1270] To implement this invention, we need to build a system that efficiently manages and shares photos and videos taken by users. This system consists of major components including smartphones, cloud storage, AI algorithms, emotion engines, and ad distribution platforms.

[1271] System program and processing flow

[1272] 1. Data upload

[1273] Users use a smartphone to upload photos and videos taken using a dedicated application to the server. The uploaded data is stored in cloud storage. During this process, users are provided with a simple interface and can upload multiple files at once.

[1274] 2. Metadata Acquisition and Storage

[1275] Metadata (e.g., shooting date and time, location, device information) is extracted from photos and videos stored in cloud storage and stored in a database. Image processing software (e.g., Google Cloud Vision) is used to obtain the metadata.

[1276] 3. Data selection using AI

[1277] An AI algorithm (e.g., OpenAI GPT-4) analyzes photos and videos from a specified period based on criteria, such as image quality, composition, facial expression, lighting, and the presence or absence of blur or noise, to select high-quality data.

[1278] 4. Emotion Recognition and Data Modification

[1279] An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in photos and videos to obtain emotional data. This emotional data is then reflected in the selection process. Furthermore, generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[1280] 5. Album creation and sharing

[1281] The server automatically compiles the edited photos and videos into an album, which can then be shared with designated recipients (e.g., grandparents) via a link sent to their email address or a notification sent to a dedicated application.

[1282] 6. Ad Targeting and Delivery

[1283] Generate advertising targets based on stored data, user profiles, and emotional data. Display personalized ads to the generated advertising targets. This is done using an advertising distribution platform (e.g., Google AdSense).

[1284] Examples of concrete examples and prompts

[1285] As a concrete example, a dual-income household user uploads photos and videos taken at a park on a day off through a dedicated application. The server stores these data in cloud storage, retrieves metadata, and organizes them. AI and an emotion engine select the best photos of the month and correct unwanted background parts. The system then creates an album using the corrected photos and videos, and automatically shares the link with grandparents. Based on the user's profile and emotion data, personalized travel-related advertisements are displayed.

[1286] Examples of prompts:

[1287] Analyze the metadata of photos and videos uploaded by user A over the past month, and use the emotional data obtained from the emotion engine to generate the most appropriate advertising targets.

[1288] Based on the user's emotional data and activity history, create and deliver ads that are likely to be of interest to them. For example, if there are a lot of travel photos and videos, generate travel-related ads, or if there are a lot of photos of sporting events, generate ads for sports equipment.

[1289] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1290] Step 1:

[1291] Users use their smartphones to upload the photos and videos they have taken to the server via a dedicated application.

[1292] Input: Photo and video files taken with a smartphone.

[1293] Output: Photo and video files saved in cloud storage.

[1294] What it does: A user opens the app, selects a photo or video to upload, and the app makes a request to send the file to the server, which receives it and stores it in cloud storage (e.g., AWS S3, Google Cloud Storage).

[1295] Step 2:

[1296] The server extracts metadata (e.g., shooting date and time, location, device information) from photos and videos stored in cloud storage and stores it in a database.

[1297] Input: Photo and video files stored in cloud storage.

[1298] Output: Metadata stored in a database.

[1299] What it does: The server uses image processing software (e.g., Google Cloud Vision) to analyze and extract metadata from photos and videos and store it in a database.

[1300] Step 3:

[1301] The server analyzes photos and videos from a specified period based on evaluation criteria and selects high-quality data.

[1302] Input: Metadata stored in the database and photo and video files.

[1303] Output: A list of selected high-quality photos and videos.

[1304] How it works: The server uses an AI algorithm (e.g., OpenAI GPT-4) to analyze photos and videos from a specified period (e.g., the past month) based on evaluation criteria (image quality, composition, facial expression, lighting conditions, presence or absence of blur and noise), and selects the most suitable ones.

[1305] Step 4:

[1306] The server uses an emotion engine to obtain emotion data from selected photos and videos and corrects the data.

[1307] Input: Selected photos and videos, and user profile information.

[1308] Output: The retouched photo or video files.

[1309] How it works: An emotion engine (e.g., Microsoft's Azure Emotion API) recognizes facial expressions in selected photos and videos to obtain emotional data. Then, a generative AI (e.g., DALL-E) automatically corrects unwanted backgrounds and reflections in the selected photos and videos.

[1310] Step 5:

[1311] The server organizes the edited photos and videos into an album and automatically shares them with the designated recipients.

[1312] Input: The modified photo or video file.

[1313] Output: Generated album and sharing link.

[1314] What it does: The server creates an album based on the modified photos and videos, and sends a sharing link to the email address of the designated recipient (e.g., grandparents), or sends a notification to a dedicated application.

[1315] Step 6:

[1316] The server generates advertising targets based on the stored data, user profiles, and emotional data, and displays personalized advertisements.

[1317] Input: Modified photos and videos, user emotion data, and user profile data.

[1318] Output: Personalized ads.

[1319] Specific operation: The server generates advertising targets based on the user profile and emotional data, and displays personalized ads using an advertising distribution platform (e.g., Google AdSense).

[1320] This series of processes allows dual-income households to efficiently organize and share photos and videos, and enables companies to effectively deliver personalized advertisements.

[1321] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1322] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1323] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1324] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1325] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1326] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1327] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1328] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1329] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1330] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1331] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1332] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1333] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1334] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1335] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1336] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1337] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1338] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1339] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1340] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1341] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1342] The following is further disclosed regarding the above embodiment.

[1343] (Claim 1)

[1344] A way to upload photos and videos taken,

[1345] A means to retrieve and store metadata for uploaded photos and videos;

[1346] A means for selecting photos and videos within a specified period based on evaluation criteria;

[1347] A means to automatically correct unwanted backgrounds and reflections in selected photos and videos,

[1348] Edited photos and videos can be organized into albums and automatically shared with designated recipients.

[1349] A system including:

[1350] (Claim 2)

[1351] The system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[1352] (Claim 3)

[1353] 10. The system of claim 1, further comprising means for generating targeting data for the advertisement based on attributes of the recipient.

[1354] "Example 1"

[1355] (Claim 1)

[1356] A way to upload photos and videos taken,

[1357] A means to retrieve and store metadata for uploaded photos and videos;

[1358] A means to store the selected photos and videos in cloud storage and store the metadata in a database;

[1359] A means for analyzing and selecting photos and videos within a specified period based on evaluation criteria;

[1360] A method for automatically correcting unwanted backgrounds and reflections in photos and videos selected based on evaluation criteria;

[1361] Edited photos and videos can be organized into albums and automatically shared with designated recipients.

[1362] A means to send the album's shared link to the recipient by email, and a means to send a notification to a dedicated application,

[1363] A system including:

[1364] (Claim 2)

[1365] The system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[1366] (Claim 3)

[1367] 10. The system of claim 1, further comprising means for generating targeting data for the advertisement based on attributes of the recipient.

[1368] "Application Example 1"

[1369] (Claim 1)

[1370] A way to upload photos and videos taken,

[1371] A means to retrieve and store metadata for uploaded photos and videos;

[1372] A means for selecting photos and videos within a specified period based on evaluation criteria;

[1373] A means to automatically correct unwanted backgrounds and reflections in selected photos and videos,

[1374] Edited photos and videos can be organized into albums and automatically shared with designated recipients.

[1375] means for analyzing designated media data to estimate user interests;

[1376] A means for targeting and delivering advertisements based on the inference results;

[1377] A system including:

[1378] (Claim 2)

[1379] The system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[1380] (Claim 3)

[1381] 2. The system according to claim 1, wherein the system estimates a user's interests and concerns based on the content of media data captured by the user.

[1382] "Example 2: Combining Emotion Engines"

[1383] (Claim 1)

[1384] A way to upload photos and videos taken,

[1385] A method to store uploaded photos and videos in cloud storage, acquire metadata, and store it in a database.

[1386] A method to select photos and videos within a specified period based on evaluation criteria using AI algorithms and emotion engines, and

[1387] A means to automatically correct unwanted backgrounds and reflections in selected photos and videos using generative AI,

[1388] Edited photos and videos can be organized into albums and automatically shared with designated recipients via email or notifications.

[1389] An information processing system including:

[1390] (Claim 2)

[1391] 2. The information processing system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur and noise.

[1392] (Claim 3)

[1393] 10. The information processing system of claim 1, further comprising means for generating targeting data for advertising based on attributes of the recipient.

[1394] "Application example 2 when combining emotion engines"

[1395] (Claim 1)

[1396] A way to upload photos and videos taken,

[1397] A means to retrieve and store metadata for uploaded photos and videos;

[1398] A means for selecting photos and videos within a specified period based on evaluation criteria;

[1399] A means to automatically correct unwanted backgrounds and reflections in selected photos and videos,

[1400] Edited photos and videos can be organized into albums and automatically shared with designated recipients.

[1401] means for generating advertising targets based on the stored data and user profiles and displaying personalized advertisements using emotional data;

[1402] A system including:

[1403] (Claim 2)

[1404] The system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur or noise.

[1405] (Claim 3)

[1406] 10. The system of claim 1, further comprising means for generating targeting data for the advertisement based on the recipient's attributes and sentiment data. [Explanation of symbols]

[1407] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A way to upload photos and videos taken, A means to retrieve and store metadata for uploaded photos and videos; A means for selecting photos and videos within a specified period based on evaluation criteria; A means to automatically correct unwanted backgrounds and reflections in selected photos and videos, Edited photos and videos can be organized into albums and automatically shared with designated recipients. A system including:

2. The system according to claim 1, wherein the evaluation criteria include facial expression, composition, lighting conditions, and the presence or absence of blur and noise.

3. The system of claim 1 , further comprising: means for generating targeting data for the advertisement based on attributes of the recipient.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A