system

The system analyzes user data to create emotionally rich video stories, addressing the challenge of unorganized digital memories and enhancing family bonds through personalized video narratives.

JP2026073336APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Users accumulate vast amounts of unorganized photo and video data, which are not effectively utilized to strengthen family bonds and share memories.

Method used

A system that analyzes user-saved photo and video data to infer emotions, generates music and narration matching these emotions, and incorporates relevant context to create emotionally rich and personalized video stories.

Benefits of technology

Enables users to relive memories through immersive family stories, enhancing family bonds by organizing and personalizing digital data into engaging video narratives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073336000001_ABST
    Figure 2026073336000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for acquiring photographic and video data, A means for analyzing the emotions of the subject based on the aforementioned data, Means for generating music and narration that match the emotions of the subject, A means of incorporating relevant context by utilizing metadata collected during data acquisition, A means for editing and generating an original video by combining the aforementioned analysis results and generated content, A means of providing the aforementioned video to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] With the spread of general digital devices, users store a huge amount of photo and video data on a daily basis. However, these data are often not organized and are not fully utilized as tools for recollecting memories. As a result, experiences and records among family members are not effectively shared, and the family bond may become diluted. To address such problems, it is necessary to effectively organize the user's stored data and provide it as a touching story to reconfirm and strengthen the family bond.

Means for Solving the Problems

[0005] This invention provides a technology that acquires user-saved photo and video data and analyzes the emotions of the subjects. Furthermore, it edits and generates original videos by using generation techniques that create music and narration that match the emotions of the subjects and incorporate relevant context based on metadata collected at the time of data collection. This allows users to effectively reminisce about past memories and provides a new way to strengthen family bonds. With such a system, users can obtain moving and immersive family stories through digital data.

[0006] "Photo and video data" refers to a collection of still and moving images recorded by a user using a digital device.

[0007] "The subject's emotions" refers to the emotional state that can be inferred from the facial expressions and posture of the person in the photograph or video.

[0008] "Music and narration" refers to musical elements and narration created to accompany video content.

[0009] "Metadata" refers to supplementary information that is automatically recorded when photos and videos are created, and includes the date and time of shooting and location information.

[0010] "Related context" refers to the narrative or meaning constructed based on the background and surrounding information of the data.

[0011] "Original video" refers to a video work newly edited and generated by combining collected data and generated content.

[0012] A "user" refers to an entity that uses this system to organize and edit its own digital data. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, a tagged communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system that utilizes photo and video data saved by users using digital devices and automatically generates emotionally moving video works based on them. This system mainly consists of a series of processes performed by a server.

[0035] First, the server organizes the photo and video data retrieved from the user's digital storage. This includes verifying and recording metadata such as the date and time of shooting and location information.

[0036] Next, the server uses an AI model to analyze the emotions of subjects in photos and videos based on the acquired data. This analysis is then used to provide a means for generating music appropriate for each scene in the video.

[0037] Subsequently, the server automatically generates music and narration tailored to the user's personality based on the information obtained from the emotion analysis. For the narration, an AI that has learned from the user's existing data is able to mimic the user's voice to tell the story. This allows family memories to be expressed from a more individual perspective.

[0038] Furthermore, the server incorporates relevant context based on the date and location information of the footage, for example, by including music or news that was popular at the time, providing a rich narrative that reflects the historical context.

[0039] Ultimately, the server edits the generated music, narration, and context to create an original video. This video can offer users a new perspective and visually reconstruct the family story.

[0040] The device notifies the user of edited videos sent from the server, and the user can watch and download them on their own device to relive memories with family. For example, if the video captures a family trip, the emotions captured in the photos and videos taken during the trip will be analyzed and presented as a travelogue narrated by the user, along with music that was popular at the travel destination.

[0041] The above is an example of an embodiment for carrying out the present invention. This system allows users to easily enjoy high-quality family history from their saved data.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The server periodically retrieves photo and video data from the user's device or cloud storage. Along with this data, it also collects metadata such as the date and time of capture and location information. The collected data is stored in temporary storage and prepared for analysis.

[0045] Step 2:

[0046] The server inputs the acquired photo and video data into an AI model to analyze the emotions of the subjects. This model analyzes facial expressions and actions within images and videos to generate emotion scores (e.g., joy, sadness, surprise) for each image or video. The analysis results are recorded in a database and used during video editing.

[0047] Step 3:

[0048] The server generates music appropriate for each scene in the video based on the results of emotion analysis. In this music generation process, the tempo and mood of the music are adjusted according to the emotion score, based on pre-prepared music templates. At the same time, an AI narrator learns from the user's past voice data and creates narration that mimics the user's voice.

[0049] Step 4:

[0050] The server utilizes metadata collected during data collection to investigate the historical context, such as popular music and news at the time of filming. Based on this information, it prepares materials to give the video a historical background and narrative. This background context information is used during the final video editing process.

[0051] Step 5:

[0052] The server uses editing software to combine the acquired photos and videos with analysis results and generated content to create the original video. Each media file is placed on a timeline according to the date and time of shooting, and emotionally appropriate background music and narration are inserted. This creates a consistent storyline.

[0053] Step 6:

[0054] The device notifies the user that the completed video is ready in their account. The user can preview the video through the application and play it by selecting streaming or download options. The user then has the opportunity to watch the generated video with family and reminisce about memories.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] In modern times, many people generate vast amounts of photographic and video data using digital devices. However, meaningfully utilizing this data and reconstructing it into emotionally impactful visual works is a significant challenge. Furthermore, there is a need for methods that can effectively express the emotions embedded in individual data and the cultural context of the time period.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes a mechanism for acquiring image and video information, means for evaluating the emotions of the subject, and technology for generating sounds and narration that are appropriate to the evaluation results. This makes it possible to automatically generate emotionally rich video works from the user's digital data and to visually and emotionally reconstruct individual memories.

[0060] "Image and video information" refers to visual media data such as photographs and videos acquired by digital devices.

[0061] "Evaluating the emotions of a subject" refers to the process of analyzing the facial expressions and atmosphere of a subject in a photograph or video and quantifying their emotional state.

[0062] "Technology for generating sound and narration" refers to technology that automatically creates suitable music and narration based on the results of an emotional evaluation of a subject.

[0063] A "mechanism" refers to a system or device designed to perform a specific function.

[0064] "Auxiliary information" refers to metadata such as date, time, and location information that accompanies data collection.

[0065] "Relevant background information" refers to historical and cultural data used to supplement the context of captured images and videos.

[0066] "Editing and creation methods" refers to the entire process of combining acquired data with generated content to construct a new video work.

[0067] "Original video" refers to video works that are individually produced to express a specific user experience or story.

[0068] This invention is a system that utilizes image and video information saved by users using digital devices and automatically generates emotionally impactful video works based on them. The system is primarily processed on a server and consists of the following hardware and software.

[0069] The server retrieves image and video data from the user's digital device or cloud storage. To achieve this, it utilizes cloud storage APIs and has the capability to access data and download necessary files. Next, the server organizes the retrieved data along with metadata for the photos and videos. A database system is used for analyzing the metadata, recording the date and time of capture and location information.

[0070] The server uses an AI model to analyze the emotions of the subject. It incorporates a process that utilizes image recognition technology to identify the subject and quantify their emotional state. The results of the emotional analysis are tagged and used to generate music and narration.

[0071] The server then uses a generative AI model to create sounds tailored to the user's personality. The music generation employs algorithms to generate melodies that match emotions, and by mimicking the user's voice using speech synthesis technology, it enables narration with a more personalized perspective.

[0072] Furthermore, the server incorporates relevant background information based on the date and time of shooting and location data, for example, by reflecting music and news that were popular at the time the data was collected, thereby providing a richer narrative. This video editing blends multiple elements of music, narration, and emotionally rich background information.

[0073] The device receives the edited video sent from the server and notifies the user. The user can then watch or download the video via this notification and relive personal memories.

[0074] As a concrete example, in a video summarizing a family trip, the system analyzes the emotions captured in photos and videos taken during the trip and generates a travelogue narrated with appropriate background music and the user's voice. This entire process can be initiated by entering a prompt such as, "Please analyze the emotions from photos and videos of our family trip, add popular music from the destination, and generate an original video with narration in my voice."

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The server retrieves image and video data from the user's digital device or cloud storage. The input for this step is the path or URL of the data specified by the user. The server accesses the data using cloud storage APIs and downloads the files. As output, the retrieved photos and videos are stored in the server's temporary storage.

[0078] Step 2:

[0079] The server analyzes the metadata of the acquired images and videos and stores it in a database. The input is the media files acquired in step 1. The server extracts the date and time of capture and location information from these files and records it in the database. The output is organized image and video data with metadata.

[0080] Step 3:

[0081] The server uses image recognition technology to evaluate the emotions of the subject. The input for this step is the metadata-enhanced image and video data organized in step 2. The server analyzes the subject's facial expressions and atmosphere via an AI model and quantifies their emotional state. The output is a dataset with emotion tags assigned to each media file.

[0082] Step 4:

[0083] The server generates sound and narration based on emotion tags using a generative AI model. The input is a dataset to which emotion tags have been assigned in step 3. The server generates melodies that match the emotion using an application program and creates narration that mimics the user's voice using speech synthesis technology. The output is a music file and a narration file.

[0084] Step 5:

[0085] The server performs video editing, incorporating relevant background information based on the time period in which the digital data was collected. The inputs for this step are the audio and narration generated in step 4, and the metadata organized in step 2. The server evaluates and incorporates music and news that were popular during the collection period. The output is the edited original video file, with all elements integrated.

[0086] Step 6:

[0087] The device notifies the user that the edited video sent from the server is complete. The input is the edited video data, which is the output of step 5. The device uses a notification API to inform the user that the video is complete and provides a download link. The user can watch the video via this notification and relive their memories.

[0088] (Application Example 1)

[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0090] In today's world, transforming the vast amount of photos and videos users capture daily into emotionally resonant visual experiences requires considerable time and effort. Furthermore, there is a lack of ways to enjoy this generated content in a more immersive form, not only in traditional two-dimensional formats, but also through augmented reality. Solving these challenges will enable users to consume their digital content as a richer experience.

[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0092] In this invention, the server includes means for acquiring photographic and video information, means for analyzing the emotions of a subject based on the information, means for generating music and audio commentary that matches the emotions of the subject, means for incorporating the relevant environment using metadata collected during information gathering, means for editing and generating original video by combining the analysis results and generated information, and means for providing the generated video in augmented reality format. As a result, users can visually experience emotionally moving videos automatically generated based on their own photos and videos, and further gain deeper emotion and understanding through augmented reality.

[0093] "Photographic and video information" refers to still images and video data acquired using digital devices.

[0094] "Subject's emotions" refers to information that indicates a specific emotional state, obtained by analyzing the emotions of people, animals, etc., captured in photographs and videos using an AI model.

[0095] "Music and audio commentary" refers to a combination of sounds and narration that match the emotions of the subject and the related story.

[0096] "Metadata" refers to data that includes supplementary information such as the date and time a photo or video was taken, location information, and camera settings.

[0097] "Relevant environment" refers to contextual information used to incorporate the cultural, historical, and geographical background of the time of filming into the video work.

[0098] "Original video" refers to a video work that has been independently edited and generated based on the analysis results and generated content.

[0099] "Augmented reality" refers to a technological format that overlays digital content onto the real world.

[0100] The system for implementing this invention begins by transmitting photographic and video information acquired by the user's digital device to a server. The server uses this information to perform analysis and detect and classify the emotions of the subjects. An AI model is used for emotion analysis, performing calculations to identify emotion patterns. This process yields emotional data from still images and videos.

[0101] Next, the server automatically generates music and audio commentary that matches the subject based on the acquired emotional data. This process uses a generative AI model to create music segments and audio narration that correspond to the analyzed emotional state.

[0102] Furthermore, the server uses metadata associated with the information collected, such as the date and time of shooting and location information, to build a related environment. This related environment includes the cultural background and geographical elements of that period, allowing the video to reflect a sense of time and place. Based on this, a unique video for the user is generated.

[0103] The generated video is then provided to the user's device in augmented reality format. The device displays this augmented reality video, providing the user with a more immersive experience.

[0104] As a concrete example, consider a scenario where a user uploads photos of their child's sports day to a server. The server performs sentiment analysis on these photos, detecting energetic and lively emotions. Based on this, it generates upbeat music and creates voice narration suitable for parental encouragement. Furthermore, it incorporates weather information for the day of the sports day and the cultural background of the specific location as relevant environmental elements into the video. Finally, the generated video is delivered in augmented reality format through a smartphone application. This is achieved by the prompt statement, "Generate an energetic and lively video based on photos of a child's sports day, and incorporate background information specific to that day."

[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0106] Step 1:

[0107] Users upload photos and videos taken with their digital devices to the server. The input consists of the user's digital media files. The server receives this data and prepares it for analysis.

[0108] Step 2:

[0109] The server extracts metadata from uploaded photos and videos. This metadata includes the date and time of capture and location information. This input metadata is used later to build the associated environment. The server records this information in a database.

[0110] Step 3:

[0111] The server uses a generative AI model to analyze the emotions of subjects in photos and videos. Image data is used as input, and the output is estimated emotion information of the subjects. The generative AI model used in this process is trained using TENSORFLOW® or PyTorch.

[0112] Step 4:

[0113] The server generates music and audio commentary that match the estimated emotions based on the results of emotion analysis. The input is analyzed emotion information, and the generating AI model outputs music and audio narration. The audio narration is generated by the AI ​​mimicking the user's voice.

[0114] Step 5:

[0115] The server constructs a related environment based on the extracted metadata, corresponding to the date and time of shooting and location information. This related environment includes cultural or geographical background information and provides information for incorporation into the video. Shooting information is used as input, and related environment data is generated as output.

[0116] Step 6:

[0117] The server combines the acquired emotion analysis results, generated music and audio commentary, and relevant environmental data to edit and produce original video. All intermediate products are inputs, and the final video is provided as output.

[0118] Step 7:

[0119] The device notifies the user of the final video content sent from the server. The user receives this video, which is provided in augmented reality format, and by experiencing the video displayed on the device, they can enjoy a more immersive experience of their memories.

[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0121] This invention is a system that uses an emotion engine to recognize the user's emotions based on user-saved photo and video data, and highly personalizes the generation of video content. This system functions primarily through a series of processes performed on a server and user interaction.

[0122] First, the server securely retrieves photo and video data from the user's digital storage and aggregates metadata, including the date and time of capture and location information. This data is temporarily stored on the server in preparation for the next processing step.

[0123] Next, the server uses an AI model to analyze the emotions of the subjects in the acquired photos and videos. In addition, it utilizes an emotion engine to collect the user's voice, facial expressions, gaze, etc., in real time through sensors, and incorporates the user's own emotional state into the system.

[0124] Furthermore, the server generates music and narration appropriate to each scene of the video based on emotion analysis and data derived from user emotions. In this process, AI technology is used to mimic the user's voice in the narration, resulting in a more relatable narrative. Additionally, based on metadata collected during data collection, the system incorporates context into the video, considering the historical background and trends of the time of filming.

[0125] Here, the server aggregates this information, drives the editing software, and generates the final original video. It can also adjust the tempo and atmosphere of the video as needed, reflecting the results of real-time sentiment analysis of the user.

[0126] Ultimately, the device notifies the user of the generated video, and the user views the video on their device. This video incorporates the user's emotions and historical context, providing a new way to revisit family memories. For example, when the user expresses joy, the mood of the video is brightened, and the narration is adjusted to reflect that emotion. In this way, the present invention provides advanced personalization that takes the user's emotions into consideration, supporting the creation of useful video content that deepens family bonds.

[0127] The following describes the processing flow.

[0128] Step 1:

[0129] The server periodically collects photos and videos stored by users in their digital storage. During this process, metadata such as the date and time of capture and location information is also retrieved. The data is stored in temporary storage for analysis, with security considerations in mind.

[0130] Step 2:

[0131] The server analyzes the acquired photo and video data using an emotion engine. This process calculates an emotion score from the facial expressions and actions of subjects in the images and videos, and further infers and records the user's emotions using voice input from the user and audio, facial, and gaze data collected in real time through the device's camera.

[0132] Step 3:

[0133] The server generates music and AI narration appropriate for the video scene based on the analyzed emotion data. The narration is generated using AI technology that has learned from the user's past voice data, mimicking the user's voice. The music is created by adjusting the tempo and key to match the analyzed emotions.

[0134] Step 4:

[0135] The server uses aggregated metadata and user sentiment information to reflect historical data and trends from the time of filming in the video. This information influences the selection of narration and music within the video and is used to provide overall context.

[0136] Step 5:

[0137] The server uses editing software to edit photo and video data in conjunction with sentiment analysis and generated content. Generated music and narration are appropriately inserted into video scenes to create emotionally rich videos. Here, the tempo and mood of the video are dynamically adjusted based on the user's real-time sentiment data.

[0138] Step 6:

[0139] The device notifies the user when the generated video is ready. The user can watch the video through the application and choose between streaming or downloading. During viewing, the user can experience the inclusion of historical data based on past events and emotionally nuanced elements.

[0140] (Example 2)

[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0142] With the widespread use of digital photos and videos, there is a challenge in how to individually optimize the vast amount of media data accumulated by users and provide a deeper emotional experience. Furthermore, existing systems have a problem in that they do not adequately generate personalized content that reflects the user's emotional state in real time.

[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0144] In this invention, the server includes a device for acquiring photographic and video data, a device for analyzing the emotions of a subject based on the data, and a device for generating music and narration based on the emotions of the subject and the emotional state of the user. This makes it possible to generate highly personalized video content that matches the user's emotions.

[0145] A "device for acquiring photo and video data" is a device for securely transferring photos and videos from a user's digital storage.

[0146] A "device for analyzing the emotions of subjects" is a device that uses an AI model to detect and analyze the emotions of subjects appearing in photographs and videos.

[0147] A "device that generates music and narration based on the user's emotional state" is a device that generates personalized music and narration based on user emotional data acquired in real time.

[0148] A "device that incorporates related information by utilizing metadata acquired during data acquisition" is a device that analyzes metadata of photos and videos (for example, date and time of shooting and location information) and incorporates the associated related information into the video content.

[0149] A "device for integrating analysis results and generated content to edit and generate personalized videos" is a device that uses the results of sentiment analysis to edit content, including music and narration, into a video, providing a personalized viewing experience for each user.

[0150] A "device that adjusts based on real-time emotional state" is a device that adjusts the tempo and atmosphere of a video by reflecting the user's current emotions during video generation.

[0151] A "device that provides and makes accessible to users" is a device that notifies users of the generated video and allows them to easily access and watch it on their own devices.

[0152] This invention is a system that generates highly personalized video content based on photos and videos owned by the user, while taking into account their emotional state. The system functions primarily through interaction between a server, a terminal, and the user. The respective components and specific embodiments are described below.

[0153] The server is the primary device for retrieving photo and video data from the user's digital storage. Network communication protocols are used for secure data retrieval. The retrieved data, along with metadata (e.g., date and location information), is temporarily stored on the server. Based on this data, an AI model is used to analyze the emotions of the subjects. The AI ​​model employs image recognition technology. Specifically, it identifies emotions such as smiles and sadness based on the subject's facial expressions.

[0154] Furthermore, the server collects the user's voice, facial expressions, and gaze in real time through sensors (camera, microphone, etc.) installed in the user's device. This activates the emotion engine, integrating the user's real-time emotional state into the system. Based on this information, AI technology is used to generate music and narration that mimics the user's voice. The generated music and narration are designed to harmonize with the user's emotions at that moment.

[0155] When generating narration, a prompt like the following might be used: "Based on the sentiment analysis results of the photos and videos, generate narration that reflects the user's feelings of joy. The narration should match the tone of the user's voice and have a cheerful mood."

[0156] Furthermore, the server integrates these analysis results and content, and uses editing software to generate the final video. During this process, historical data, built based on shooting date and location information, is used to enhance the scenes, providing richer context to the video.

[0157] Ultimately, the device notifies the user of the generated video, which the user can then freely view on their device. The video, which reflects the user's emotions and surrounding environment, offers a new experience for revisiting memories with family and loved ones.

[0158] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0159] Step 1:

[0160] The server obtains permission from the user to connect to digital storage and retrieves photo and video data. The input is the user's storage path or cloud storage access information. The retrieved data is stored on the server along with metadata such as the date and time of capture and location information. During this process, encryption protocols such as SSL are used to ensure security during data transfer.

[0161] Step 2:

[0162] The server inputs collected photo and video data into an AI model to analyze the emotions of the subjects. The input is image and video data, and the output is data indicating emotional characteristics such as smiles, sadness, and surprise. Specifically, it uses image recognition technology to identify facial features and machine learning algorithms to infer specific emotions.

[0163] Step 3:

[0164] The server collects voice, gaze, and facial expressions in real time from sensors via the user's device. The input is raw data from the microphone and camera, which is used to infer the user's current emotional state. The output is emotional data, such as whether the user is happy or relaxed. An emotion engine is used to process this information, and the analysis results are instantly reflected on the server.

[0165] Step 4:

[0166] The server generates music and narration based on the results of sentiment analysis and the user's emotional state. The input consists of sentiment data and prompt text perceived by the program. The output is music and narration with a tone and content that matches the user's emotions. Using a generative AI model, the narration is created in a natural tone based on the user's past voice data.

[0167] Step 5:

[0168] The server integrates analysis results and generated content, and uses editing software to produce personalized videos. Inputs include music, narration, and original image / video data, while output is the completed original video. This process incorporates historical context and trends based on metadata from the time of filming, striving for visual aesthetics appropriate to the content.

[0169] Step 6:

[0170] The device notifies the user of the generated video and provides it in a playable format. The input is video data from the server, and the output is video content played on the device's screen. Users can easily play this video on their own devices and enjoy a personalized viewing experience.

[0171] (Application Example 2)

[0172] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0173] When generating personalized video content from digital data, it is necessary to go beyond simply relying on static sentiment analysis and instead reflect the user's real-time emotional state to provide more relatable and empathetic content. Furthermore, it is essential to provide more personal and approachable narration to enhance the memorable experience.

[0174] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0175] In this invention, the server includes means for acquiring digital information, means for analyzing the emotions of a target based on the information, and means for collecting the user's emotional state in real time and adjusting the tempo and atmosphere in the video. This makes it possible to generate video that more faithfully responds to the user's emotions.

[0176] "Digital information" refers to visual media data stored electronically, such as photographs and videos.

[0177] "Subject" is a term that refers to a person or object that appears in a photograph or video.

[0178] "Analyzing emotions" means estimating a person's emotional state from their facial expressions and gestures in visual media.

[0179] "Musical material and explanatory text" refers to background music suitable for the video and text generated as narration, which are added to the visual content.

[0180] "Auxiliary information" refers to metadata such as location information and date / time information that is collected when taking photos or videos.

[0181] "Background" refers to the historical context and cultural background that are incorporated to make the video content easier to understand.

[0182] "Unique video content" refers to personalized video content generated based on the user's specific emotions.

[0183] "User" is a term that refers to an individual or organization using a digital information provision system.

[0184] "Real-time collection" refers to the process by which emotional data is acquired instantly as the user uses the system.

[0185] "Adjusting tempo and atmosphere" means dynamically changing the playback speed of the video and the tone of the music to match the user's emotions.

[0186] To realize this invention, the server first acquires photo and video data from the user's digital device. The data includes metadata such as the date and time of shooting and location information, and this is securely stored as digital information in a cloud environment. The server performs sentiment analysis on the image and video data using an AI model. In this process, the subject's facial expressions and behavioral patterns are analyzed, and their emotional state is estimated. Furthermore, the tempo and atmosphere of the video are dynamically adjusted based on the user's sentiment data collected in real time.

[0187] On the software side, the server uses emotion analysis engines such as Google Cloud AI and Amazon Rekognition to process data. These engines extract information about emotions from fragmented data and derive the most appropriate personalization elements. Based on the analysis results, the server generates music and explanatory text, and a generative AI model provides user-specific narration. The narration enhances personal familiarity by mimicking the user's past voice data.

[0188] As a concrete example, consider a case where photos from a family trip are uploaded to a server. In this case, if the analyzed emotion is determined to be "joy," the server will generate music that creates a joyful atmosphere, along with narration reflecting on the feelings at the time of the trip. Once the video is complete, the device will notify the user of the generated video, making it available for viewing.

[0189] An example of a prompt for a generative AI model is, "Extract fun moments from a family photo album and create an exciting music video."

[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0191] Step 1:

[0192] Users upload photos and videos to the cloud using their smartphones. The input consists of photos and videos taken from the user's digital device. The output is that this data is stored securely on a server.

[0193] Step 2:

[0194] The server analyzes metadata (such as date and location information) of photos and videos stored in the cloud. The input is the metadata of photos and videos uploaded by the user. The output is the analyzed metadata. This data is later used to build context for related background information.

[0195] Step 3:

[0196] The server uses an AI model to analyze the emotions of subjects from photos and videos. The input is visual media data from the user. The AI ​​model analyzes facial expressions and behavioral patterns, and generates data that identifies the emotions of the subjects as output.

[0197] Step 4:

[0198] The server collects user voice and facial expression data in real time and adjusts the video tempo and atmosphere based on that data. The input is the user's real-time voice and facial expression data. The output is adjusted video settings that take the user's emotions into consideration.

[0199] Step 5:

[0200] The server generates music and explanatory text based on the analysis results. The input is the subject's emotional data, and the output is music and narration text appropriate to that emotion.

[0201] Step 6:

[0202] The server uses a generated AI model to provide user-specific narration. Inputs include the user's past voice data and narration text. Output is narration audio that closely resembles the user's voice.

[0203] Step 7:

[0204] The server combines all elements (adjusted video, music, narration) to edit and generate a unique video. Inputs are the adjusted video settings, music, and narration. The output is a fully edited, personalized video.

[0205] Step 8:

[0206] The device notifies the user of the completed video and makes it available for viewing. The input is the edited video data. As output, the user can view the video generated on their device.

[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0210] [Second Embodiment]

[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0223] This invention is a system that utilizes photographic and video data saved by users using digital devices to automatically generate emotionally impactful video works. This system consists mainly of a series of processes performed by a server.

[0224] First, the server organizes the photo and video data retrieved from the user's digital storage. This includes verifying and recording metadata such as the date and time of shooting and location information.

[0225] Next, the server uses an AI model to analyze the emotions of subjects in photos and videos based on the acquired data. This analysis is then used to provide a means for generating music appropriate for each scene in the video.

[0226] Subsequently, the server automatically generates music and narration tailored to the user's personality based on the information obtained from the emotion analysis. For the narration, an AI that has learned from the user's existing data is able to mimic the user's voice to tell the story. This allows family memories to be expressed from a more individual perspective.

[0227] Furthermore, the server incorporates relevant context based on the date and location information of the footage, for example, by including music or news that was popular at the time, providing a rich narrative that reflects the historical context.

[0228] Ultimately, the server edits the generated music, narration, and context to create an original video. This video can offer users a new perspective and visually reconstruct the family story.

[0229] The device notifies the user of edited videos sent from the server, and the user can watch and download them on their own device to relive memories with family. For example, if the video captures a family trip, the emotions captured in the photos and videos taken during the trip will be analyzed and presented as a travelogue narrated by the user, along with music that was popular at the travel destination.

[0230] The above is an example of an embodiment for carrying out the present invention. This system allows users to easily enjoy high-quality family history from their saved data.

[0231] The following describes the processing flow.

[0232] Step 1:

[0233] The server periodically retrieves photo and video data from the user's device or cloud storage. Along with this data, it also collects metadata such as the date and time of capture and location information. The collected data is stored in temporary storage and prepared for analysis.

[0234] Step 2:

[0235] The server inputs the acquired photo and video data into an AI model to analyze the emotions of the subjects. This model analyzes facial expressions and actions within images and videos to generate emotion scores (e.g., joy, sadness, surprise) for each image or video. The analysis results are recorded in a database and used during video editing.

[0236] Step 3:

[0237] The server generates music appropriate for each scene in the video based on the results of emotion analysis. In this music generation process, the tempo and mood of the music are adjusted according to the emotion score, based on pre-prepared music templates. At the same time, an AI narrator learns from the user's past voice data and creates narration that mimics the user's voice.

[0238] Step 4:

[0239] The server utilizes metadata collected during data collection to investigate the historical context, such as popular music and news at the time of filming. Based on this information, it prepares materials to give the video a historical background and narrative. This background context information is used during the final video editing process.

[0240] Step 5:

[0241] The server uses editing software to combine the acquired photos and videos with analysis results and generated content to create the original video. Each media file is placed on a timeline according to the date and time of shooting, and emotionally appropriate background music and narration are inserted. This creates a consistent storyline.

[0242] Step 6:

[0243] The device notifies the user that the completed video is ready in their account. The user can preview the video through the application and play it by selecting streaming or download options. The user then has the opportunity to watch the generated video with family and reminisce about memories.

[0244] (Example 1)

[0245] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0246] In modern times, many people generate vast amounts of photographic and video data using digital devices. However, meaningfully utilizing this data and reconstructing it into emotionally impactful visual works is a significant challenge. Furthermore, there is a need for methods that can effectively express the emotions embedded in individual data and the cultural context of the time period.

[0247] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0248] In this invention, the server includes a mechanism for acquiring image and video information, means for evaluating the emotions of the subject, and technology for generating sounds and narration that are appropriate to the evaluation results. This makes it possible to automatically generate emotionally rich video works from the user's digital data and to visually and emotionally reconstruct individual memories.

[0249] "Image and video information" refers to visual media data such as photographs and videos acquired by digital devices.

[0250] "Evaluating the emotions of a subject" refers to the process of analyzing the facial expressions and atmosphere of a subject in a photograph or video and quantifying their emotional state.

[0251] "Technology for generating sound and narration" refers to technology that automatically creates suitable music and narration based on the results of an emotional evaluation of a subject.

[0252] A "mechanism" refers to a system or device designed to perform a specific function.

[0253] "Auxiliary information" refers to metadata such as date, time, and location information that accompanies data collection.

[0254] "Relevant background information" refers to historical and cultural data used to supplement the context of captured images and videos.

[0255] "Editing and creation methods" refers to the entire process of combining acquired data with generated content to construct a new video work.

[0256] "Original video" refers to video works that are individually produced to express a specific user experience or story.

[0257] This invention is a system that utilizes image and video information saved by users using digital devices and automatically generates emotionally impactful video works based on them. The system is primarily processed on a server and consists of the following hardware and software.

[0258] The server retrieves image and video data from the user's digital device or cloud storage. To achieve this, it utilizes cloud storage APIs and has the capability to access data and download necessary files. Next, the server organizes the retrieved data along with metadata for the photos and videos. A database system is used for analyzing the metadata, recording the date and time of capture and location information.

[0259] The server uses an AI model to analyze the emotions of the subject. It incorporates a process that utilizes image recognition technology to identify the subject and quantify its emotional state. The results of the emotional analysis are tagged and used to generate music and narration.

[0260] The server then uses a generative AI model to create sounds tailored to the user's personality. The music generation employs algorithms to generate melodies that match emotions, and by mimicking the user's voice using speech synthesis technology, it enables narration with a more personalized perspective.

[0261] Furthermore, the server incorporates relevant background information based on the date and time of shooting and location data, for example, by reflecting music and news that were popular at the time the data was collected, thereby providing a richer narrative. This video editing blends multiple elements of music, narration, and emotionally rich background information.

[0262] The device receives the edited video sent from the server and notifies the user. The user can then watch or download the video via this notification and relive personal memories.

[0263] As a concrete example, in a video summarizing a family trip, the system analyzes the emotions captured in photos and videos taken during the trip and generates a travelogue narrated with appropriate background music and the user's voice. This entire process can be initiated by entering a prompt such as, "Please analyze the emotions from photos and videos of our family trip, add popular music from the destination, and generate an original video with narration in my voice."

[0264] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0265] Step 1:

[0266] The server retrieves image and video data from the user's digital device or cloud storage. The input for this step is the path or URL of the data specified by the user. The server accesses the data using cloud storage APIs and downloads the files. As output, the retrieved photos and videos are stored in the server's temporary storage.

[0267] Step 2:

[0268] The server analyzes the metadata of the acquired images and videos and stores it in a database. The input is the media files acquired in step 1. The server extracts the date and time of capture and location information from these files and records it in the database. The output is organized image and video data with metadata.

[0269] Step 3:

[0270] The server uses image recognition technology to evaluate the emotions of the subject. The input for this step is the metadata-enhanced image and video data organized in step 2. The server analyzes the subject's facial expressions and atmosphere via an AI model and quantifies their emotional state. The output is a dataset with emotion tags assigned to each media file.

[0271] Step 4:

[0272] The server generates sound and narration based on emotion tags using a generative AI model. The input is a dataset to which emotion tags have been assigned in step 3. The server generates melodies that match the emotion using an application program and creates narration that mimics the user's voice using speech synthesis technology. The output is a music file and a narration file.

[0273] Step 5:

[0274] The server performs video editing, incorporating relevant background information based on the time period in which the digital data was collected. The inputs for this step are the audio and narration generated in step 4, and the metadata organized in step 2. The server evaluates and incorporates music and news that were popular during the collection period. The output is the edited original video file, with all elements integrated.

[0275] Step 6:

[0276] The device notifies the user that the edited video sent from the server is complete. The input is the edited video data, which is the output of step 5. The device uses a notification API to inform the user that the video is complete and provides a download link. The user can watch the video via this notification and relive their memories.

[0277] (Application Example 1)

[0278] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0279] In modern times, it takes a great deal of time and effort to experience individual memories as moving video works from the vast amount of photo and video data that users routinely shoot. Moreover, there is a lack of means to enjoy the generated content not only in the conventional flat format but also in a more immersive form through augmented reality. By solving these problems, users can consume their digital content as a richer experience.

[0280] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0281] In this invention, the server includes means for acquiring photo and video information, means for analyzing the emotions of the subject based on the information, means for generating music and voice commentary that match the emotions of the subject, means for incorporating the relevant environment using the metadata at the time of information collection, means for editing and generating an original video by combining the analysis results and the generated information, and means for providing the generated video in an augmented reality format. Thereby, users can visually experience automatically generated moving videos based on their own photos and videos, and further obtain deeper emotions and understanding through augmented reality.

[0282] "Photo and video information" refers to still image and moving image data acquired using a digital device.

[0283] "The emotion of the subject" is information indicating a specific emotional state by analyzing the emotions of people, animals, etc. shown in photos and videos using an AI model.

[0284] "Music and voice commentary" refers to a combination of sounds and narration that match the emotions of the subject and the related story.

[0285] "Metadata" is data including accompanying information such as the shooting date and time of photos and videos, location information, camera settings, etc.

[0286] The "related environment" refers to the contextual information for incorporating the cultural, historical, and geographical background at the time of shooting into video works.

[0287] The "original video" refers to a video work that is uniquely edited and generated based on the analysis results and generated content.

[0288] The "augmented reality format" refers to a technical format that overlays digital content on the real environment for display.

[0289] The system for implementing this invention begins with transmitting the photo and video information acquired by the user's digital device to the server. The server performs analysis processing using this information to detect and classify the emotions of the subject. For emotion analysis, an AI model is utilized to perform operations for identifying emotion patterns. Through this process, data related to emotions is obtained from still images and videos.

[0290] Next, based on the acquired emotion data, the server automatically generates music and voiceovers that match the subject. In this process, a generation AI model is used to create music segments and voice narrations corresponding to the analyzed emotional state.

[0291] Furthermore, the server constructs the related environment using the metadata associated with the information collection, such as the shooting date and location information. Since this related environment includes the cultural background and geographical elements of that period, the era characteristics and locality can be reflected in the video. Based on this, an original video for the user is generated.

[0292] The generated video is then provided to the user's terminal in the augmented reality format. By displaying this augmented reality video, the terminal brings a highly immersive experience to the user.

[0293] As a concrete example, consider a scenario where a user uploads photos of their child's sports day to a server. The server performs sentiment analysis on these photos, detecting energetic and lively emotions. Based on this, it generates upbeat music and creates voice narration suitable for parental encouragement. Furthermore, it incorporates weather information for the day of the sports day and the cultural background of the specific location as relevant environmental elements into the video. Finally, the generated video is delivered in augmented reality format through a smartphone application. This is achieved by the prompt statement, "Generate an energetic and lively video based on photos of a child's sports day, and incorporate background information specific to that day."

[0294] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0295] Step 1:

[0296] Users upload photos and videos taken with their digital devices to the server. The input consists of the user's digital media files. The server receives this data and prepares it for analysis.

[0297] Step 2:

[0298] The server extracts metadata from uploaded photos and videos. This metadata includes the date and time of capture and location information. This input metadata is used later to build the associated environment. The server records this information in a database.

[0299] Step 3:

[0300] The server uses a generative AI model to analyze the emotions of subjects in photos and videos. Image data is used as input, and the output is estimated emotion information of the subjects. The generative AI model used in this process is trained using TensorFlow or PyTorch.

[0301] Step 4:

[0302] Based on the result of sentiment analysis, the server generates music and voice commentary that match the estimated sentiment. What is input is the analyzed sentiment information, and music and voice narration are output using a generation AI model. The voice narration is generated by the AI imitating the user's voice.

[0303] Step 5:

[0304] Based on the extracted metadata, the server constructs a related environment according to the shooting date and time and location information. This related environment includes cultural or geographical backgrounds and provides information for incorporation into the video. Shooting information is used as input, and data of the related environment is generated as output.

[0305] Step 6:

[0306] The server combines the obtained sentiment analysis results, the generated music and voice commentary, and the related environment data to edit and generate an original video. The inputs here are all the intermediate products, and a completed video work is provided as output.

[0307] Step 7:

[0308] The terminal notifies the user of the final video work sent from the server. The user can receive this video provided in the augmented reality format and experience the video projected by the terminal to enjoy a more immersive memory.

[0309] Furthermore, a sentiment engine for estimating the user's sentiment may be combined. That is, the specific processing unit 290 may estimate the user's sentiment using the sentiment specific model 59 and perform specific processing using the user's sentiment.

[0310] This invention is a system that uses an emotion engine to recognize the user's emotions based on user-saved photo and video data, and highly personalizes the generation of video content. This system functions primarily through a series of processes performed on a server and user interaction.

[0311] First, the server securely retrieves photo and video data from the user's digital storage and aggregates metadata, including the date and time of capture and location information. This data is temporarily stored on the server in preparation for the next processing step.

[0312] Next, the server uses an AI model to analyze the emotions of the subjects in the acquired photos and videos. In addition, it utilizes an emotion engine to collect the user's voice, facial expressions, gaze, etc., in real time through sensors, and incorporates the user's own emotional state into the system.

[0313] Furthermore, the server generates music and narration appropriate to each scene of the video based on emotion analysis and data derived from user emotions. In this process, AI technology is used to mimic the user's voice in the narration, resulting in a more relatable narrative. Additionally, based on metadata collected during data collection, the system incorporates context into the video, considering the historical background and trends of the time of filming.

[0314] Here, the server aggregates this information, drives the editing software, and generates the final original video. It can also adjust the tempo and atmosphere of the video as needed, reflecting the results of real-time sentiment analysis of the user.

[0315] Ultimately, the device notifies the user of the generated video, and the user views the video on their device. This video incorporates the user's emotions and historical context, providing a new way to revisit family memories. For example, when the user expresses joy, the mood of the video is brightened, and the narration is adjusted to reflect that emotion. In this way, the present invention provides advanced personalization that takes the user's emotions into consideration, supporting the creation of useful video content that deepens family bonds.

[0316] The following describes the processing flow.

[0317] Step 1:

[0318] The server periodically collects photos and videos stored by users in their digital storage. During this process, metadata such as the date and time of capture and location information is also retrieved. The data is stored in temporary storage for analysis, with security considerations in mind.

[0319] Step 2:

[0320] The server analyzes the acquired photo and video data using an emotion engine. This process calculates an emotion score from the facial expressions and actions of subjects in the images and videos, and further infers and records the user's emotions using voice input from the user and audio, facial, and gaze data collected in real time through the device's camera.

[0321] Step 3:

[0322] The server generates music and AI narration appropriate for the video scene based on the analyzed emotion data. The narration is generated using AI technology that has learned from the user's past voice data, mimicking the user's voice. The music is created by adjusting the tempo and key to match the analyzed emotions.

[0323] Step 4:

[0324] The server uses aggregated metadata and user sentiment information to reflect historical data and trends from the time of filming in the video. This information influences the selection of narration and music within the video and is used to provide overall context.

[0325] Step 5:

[0326] The server uses editing software to edit photo and video data in conjunction with sentiment analysis and generated content. Generated music and narration are appropriately inserted into video scenes to create emotionally rich videos. Here, the tempo and mood of the video are dynamically adjusted based on the user's real-time sentiment data.

[0327] Step 6:

[0328] The device notifies the user when the generated video is ready. The user can watch the video through the application and choose between streaming or downloading. During viewing, the user can experience the inclusion of historical data based on past events and emotionally nuanced elements.

[0329] (Example 2)

[0330] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0331] With the widespread use of digital photos and videos, there is a challenge in how to individually optimize the vast amount of media data accumulated by users and provide a deeper emotional experience. Furthermore, existing systems have a problem in that they do not adequately generate personalized content that reflects the user's emotional state in real time.

[0332] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0333] In this invention, the server includes a device for acquiring photographic and video data, a device for analyzing the emotions of a subject based on the data, and a device for generating music and narration based on the emotions of the subject and the emotional state of the user. This makes it possible to generate highly personalized video content that matches the user's emotions.

[0334] A "device for acquiring photo and video data" is a device for securely transferring photos and videos from a user's digital storage.

[0335] A "device for analyzing the emotions of subjects" is a device that uses an AI model to detect and analyze the emotions of subjects appearing in photographs and videos.

[0336] A "device that generates music and narration based on the user's emotional state" is a device that generates personalized music and narration based on user emotional data acquired in real time.

[0337] A "device that incorporates related information by utilizing metadata acquired during data acquisition" is a device that analyzes metadata of photos and videos (for example, date and time of shooting and location information) and incorporates the associated related information into the video content.

[0338] A "device for integrating analysis results and generated content to edit and generate personalized videos" is a device that uses the results of sentiment analysis to edit content, including music and narration, into a video, providing a personalized viewing experience for each user.

[0339] A "device that adjusts based on real-time emotional state" is a device that adjusts the tempo and atmosphere of a video by reflecting the user's current emotions during video generation.

[0340] A "device that provides and makes accessible to users" is a device that notifies users of the generated video and allows them to easily access and watch it on their own devices.

[0341] This invention is a system that generates highly personalized video content based on photos and videos owned by the user, while taking into account their emotional state. The system functions primarily through interaction between a server, a terminal, and the user. The respective components and specific embodiments are described below.

[0342] The server is the primary device for retrieving photo and video data from the user's digital storage. Network communication protocols are used for secure data retrieval. The retrieved data, along with metadata (e.g., date and location information), is temporarily stored on the server. Based on this data, an AI model is used to analyze the emotions of the subjects. The AI ​​model employs image recognition technology. Specifically, it identifies emotions such as smiles and sadness based on the subject's facial expressions.

[0343] Furthermore, the server collects the user's voice, facial expressions, and gaze in real time through sensors (camera, microphone, etc.) installed in the user's device. This activates the emotion engine, integrating the user's real-time emotional state into the system. Based on this information, AI technology is used to generate music and narration that mimics the user's voice. The generated music and narration are designed to harmonize with the user's emotions at that moment.

[0344] When generating narration, a prompt like the following might be used: "Based on the sentiment analysis results of the photos and videos, generate narration that reflects the user's feelings of joy. The narration should match the tone of the user's voice and have a cheerful mood."

[0345] Furthermore, the server integrates these analysis results and content, and uses editing software to generate the final video. During this process, historical data, built based on shooting date and location information, is used to enhance the scenes, providing richer context to the video.

[0346] Ultimately, the device notifies the user of the generated video, which the user can then freely view on their device. The video, which reflects the user's emotions and surrounding environment, offers a new experience for revisiting memories with family and loved ones.

[0347] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0348] Step 1:

[0349] The server obtains permission from the user to connect to digital storage and retrieves photo and video data. The input is the user's storage path or cloud storage access information. The retrieved data is stored on the server along with metadata such as the date and time of capture and location information. During this process, encryption protocols such as SSL are used to ensure security during data transfer.

[0350] Step 2:

[0351] The server inputs collected photo and video data into an AI model to analyze the emotions of the subjects. The input is image and video data, and the output is data indicating emotional characteristics such as smiles, sadness, and surprise. Specifically, it uses image recognition technology to identify facial features and machine learning algorithms to infer specific emotions.

[0352] Step 3:

[0353] The server collects voice, gaze, and facial expressions in real time from sensors via the user's device. The input is raw data from the microphone and camera, which is used to infer the user's current emotional state. The output is emotional data, such as whether the user is happy or relaxed. An emotion engine is used to process this information, and the analysis results are instantly reflected on the server.

[0354] Step 4:

[0355] The server generates music and narration based on the results of sentiment analysis and the user's emotional state. The input consists of sentiment data and prompt text perceived by the program. The output is music and narration with a tone and content that matches the user's emotions. Using a generative AI model, the narration is created in a natural tone based on the user's past voice data.

[0356] Step 5:

[0357] The server integrates analysis results and generated content, and uses editing software to produce personalized videos. Inputs include music, narration, and original image / video data, while output is the completed original video. This process incorporates historical context and trends based on metadata from the time of filming, striving for visual aesthetics appropriate to the content.

[0358] Step 6:

[0359] The device notifies the user of the generated video and provides it in a playable format. The input is video data from the server, and the output is video content played on the device's screen. Users can easily play this video on their own devices and enjoy a personalized viewing experience.

[0360] (Application Example 2)

[0361] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0362] When generating personalized video content from digital data, it is necessary to go beyond simply relying on static sentiment analysis and instead reflect the user's real-time emotional state to provide more relatable and empathetic content. Furthermore, it is essential to provide more personal and approachable narration to enhance the memorable experience.

[0363] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0364] In this invention, the server includes means for acquiring digital information, means for analyzing the emotions of a target based on the information, and means for collecting the user's emotional state in real time and adjusting the tempo and atmosphere in the video. This makes it possible to generate video that more faithfully responds to the user's emotions.

[0365] "Digital information" refers to visual media data stored electronically, such as photographs and videos.

[0366] "Subject" is a term that refers to a person or object that appears in a photograph or video.

[0367] "Analyzing emotions" means estimating a person's emotional state from their facial expressions and gestures in visual media.

[0368] "Musical material and explanatory text" refers to background music suitable for the video and text generated as narration, which are added to the visual content.

[0369] "Auxiliary information" refers to metadata such as location information and date / time information that is collected when taking photos or videos.

[0370] "Background" refers to the historical context and cultural background that are incorporated to make the video content easier to understand.

[0371] "Unique video content" refers to personalized video content generated based on the user's specific emotions.

[0372] "User" is a term that refers to an individual or organization using a digital information provision system.

[0373] "Real-time collection" refers to the process by which emotional data is acquired instantly as the user uses the system.

[0374] "Adjusting tempo and atmosphere" means dynamically changing the playback speed of the video and the tone of the music to match the user's emotions.

[0375] To realize this invention, the server first acquires photo and video data from the user's digital device. The data includes metadata such as the date and time of shooting and location information, and this is securely stored as digital information in a cloud environment. The server performs sentiment analysis on the image and video data using an AI model. In this process, the subject's facial expressions and behavioral patterns are analyzed, and their emotional state is estimated. Furthermore, the tempo and atmosphere of the video are dynamically adjusted based on the user's sentiment data collected in real time.

[0376] On the software side, the server uses sentiment analysis engines such as Google Cloud AI and Amazon Rekognition to process data. These engines extract emotional information from fragmented data and derive the most appropriate personalization elements. Based on the analysis results, the server generates music and explanatory text, and a generative AI model provides user-specific narration. The narration enhances personal familiarity by mimicking the user's past voice data.

[0377] As a concrete example, consider a case where photos from a family trip are uploaded to a server. In this case, if the analyzed emotion is determined to be "joy," the server will generate music that creates a joyful atmosphere, along with narration reflecting on the feelings at the time of the trip. Once the video is complete, the device will notify the user of the generated video, making it available for viewing.

[0378] An example of a prompt for a generative AI model is, "Extract fun moments from a family photo album and create an exciting music video."

[0379] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0380] Step 1:

[0381] Users upload photos and videos to the cloud using their smartphones. The input consists of photos and videos taken from the user's digital device. The output is that this data is stored securely on a server.

[0382] Step 2:

[0383] The server analyzes metadata (such as date and location information) of photos and videos stored in the cloud. The input is the metadata of photos and videos uploaded by the user. The output is the analyzed metadata. This data is later used to build context for related background information.

[0384] Step 3:

[0385] The server uses an AI model to analyze the emotions of subjects from photos and videos. The input is visual media data from the user. The AI ​​model analyzes facial expressions and behavioral patterns, and generates data that identifies the emotions of the subjects as output.

[0386] Step 4:

[0387] The server collects user voice and facial expression data in real time and adjusts the video tempo and atmosphere based on that data. The input is the user's real-time voice and facial expression data. The output is adjusted video settings that take the user's emotions into consideration.

[0388] Step 5:

[0389] The server generates music and explanatory text based on the analysis results. The input is the subject's emotional data, and the output is music and narration text appropriate to that emotion.

[0390] Step 6:

[0391] The server uses a generated AI model to provide user-specific narration. Inputs include the user's past voice data and narration text. Output is narration audio that closely resembles the user's voice.

[0392] Step 7:

[0393] The server combines all elements (adjusted video, music, narration) to edit and generate a unique video. Inputs are the adjusted video settings, music, and narration. The output is a fully edited, personalized video.

[0394] Step 8:

[0395] The device notifies the user of the completed video and makes it available for viewing. The input is the edited video data. As output, the user can view the video generated on their device.

[0396] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0397] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0398] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0399] [Third Embodiment]

[0400] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0401] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0402] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0403] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0404] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0406] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0407] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0408] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0409] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0410] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0411] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0412] This invention is a system that utilizes photographic and video data saved by users using digital devices to automatically generate emotionally impactful video works. This system consists mainly of a series of processes performed by a server.

[0413] First, the server organizes the photo and video data retrieved from the user's digital storage. This includes verifying and recording metadata such as the date and time of shooting and location information.

[0414] Next, the server uses an AI model to analyze the emotions of subjects in photos and videos based on the acquired data. This analysis is then used to provide a means for generating music appropriate for each scene in the video.

[0415] Subsequently, the server automatically generates music and narration tailored to the user's personality based on the information obtained from the emotion analysis. For the narration, an AI that has learned from the user's existing data is able to mimic the user's voice to tell the story. This allows family memories to be expressed from a more individual perspective.

[0416] Furthermore, the server incorporates relevant context based on the date and location information of the footage, for example, by including music or news that was popular at the time, providing a rich narrative that reflects the historical context.

[0417] Ultimately, the server edits the generated music, narration, and context to create an original video. This video can offer users a new perspective and visually reconstruct the family story.

[0418] The device notifies the user of edited videos sent from the server, and the user can watch and download them on their own device to relive memories with family. For example, if the video captures a family trip, the emotions captured in the photos and videos taken during the trip will be analyzed and presented as a travelogue narrated by the user, along with music that was popular at the travel destination.

[0419] The above is an example of an embodiment for carrying out the present invention. This system allows users to easily enjoy high-quality family history from their saved data.

[0420] The following describes the processing flow.

[0421] Step 1:

[0422] The server periodically retrieves photo and video data from the user's device or cloud storage. Along with this data, it also collects metadata such as the date and time of capture and location information. The collected data is stored in temporary storage and prepared for analysis.

[0423] Step 2:

[0424] The server inputs the acquired photo and video data into an AI model to analyze the emotions of the subjects. This model analyzes facial expressions and actions within images and videos to generate emotion scores (e.g., joy, sadness, surprise) for each image or video. The analysis results are recorded in a database and used during video editing.

[0425] Step 3:

[0426] The server generates music appropriate for each scene in the video based on the results of emotion analysis. In this music generation process, the tempo and mood of the music are adjusted according to the emotion score, based on pre-prepared music templates. At the same time, an AI narrator learns from the user's past voice data and creates narration that mimics the user's voice.

[0427] Step 4:

[0428] The server utilizes metadata collected during data collection to investigate the historical context, such as popular music and news at the time of filming. Based on this information, it prepares materials to give the video a historical background and narrative. This background context information is used during the final video editing process.

[0429] Step 5:

[0430] The server uses editing software to combine the acquired photos and videos with analysis results and generated content to create the original video. Each media file is placed on a timeline according to the date and time of shooting, and emotionally appropriate background music and narration are inserted. This creates a consistent storyline.

[0431] Step 6:

[0432] The device notifies the user that the completed video is ready in their account. The user can preview the video through the application and play it by selecting streaming or download options. The user then has the opportunity to watch the generated video with family and reminisce about memories.

[0433] (Example 1)

[0434] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0435] In modern times, many people generate vast amounts of photographic and video data using digital devices. However, meaningfully utilizing this data and reconstructing it into emotionally impactful visual works is a significant challenge. Furthermore, there is a need for methods that can effectively express the emotions embedded in individual data and the cultural context of the time period.

[0436] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0437] In this invention, the server includes a mechanism for acquiring image and video information, means for evaluating the emotions of the subject, and technology for generating sounds and narration that are appropriate to the evaluation results. This makes it possible to automatically generate emotionally rich video works from the user's digital data and to visually and emotionally reconstruct individual memories.

[0438] "Image and video information" refers to visual media data such as photographs and videos acquired by digital devices.

[0439] "Evaluating the emotions of a subject" refers to the process of analyzing the facial expressions and atmosphere of a subject in a photograph or video and quantifying their emotional state.

[0440] "Technology for generating sound and narration" refers to technology that automatically creates suitable music and narration based on the results of an emotional evaluation of a subject.

[0441] A "mechanism" refers to a system or device designed to perform a specific function.

[0442] "Auxiliary information" refers to metadata such as date, time, and location information that accompanies data collection.

[0443] "Relevant background information" refers to historical and cultural data used to supplement the context of captured images and videos.

[0444] "Editing and creation methods" refers to the entire process of combining acquired data with generated content to construct a new video work.

[0445] "Original video" refers to video works that are individually produced to express a specific user experience or story.

[0446] This invention is a system that utilizes image and video information saved by users using digital devices and automatically generates emotionally impactful video works based on them. The system is primarily processed on a server and consists of the following hardware and software.

[0447] The server retrieves image and video data from the user's digital device or cloud storage. To achieve this, it utilizes cloud storage APIs and has the capability to access data and download necessary files. Next, the server organizes the retrieved data along with metadata for the photos and videos. A database system is used for analyzing the metadata, recording the date and time of capture and location information.

[0448] The server uses an AI model to analyze the emotions of the subject. It incorporates a process that utilizes image recognition technology to identify the subject and quantify its emotional state. The results of the emotional analysis are tagged and used to generate music and narration.

[0449] The server then uses a generative AI model to create sounds tailored to the user's personality. The music generation employs algorithms to generate melodies that match emotions, and by mimicking the user's voice using speech synthesis technology, it enables narration with a more personalized perspective.

[0450] Furthermore, the server incorporates relevant background information based on the date and time of shooting and location data, for example, by reflecting music and news that were popular at the time the data was collected, thereby providing a richer narrative. This video editing blends multiple elements of music, narration, and emotionally rich background information.

[0451] The device receives the edited video sent from the server and notifies the user. The user can then watch or download the video via this notification and relive personal memories.

[0452] As a concrete example, in a video summarizing a family trip, the system analyzes the emotions captured in photos and videos taken during the trip and generates a travelogue narrated with appropriate background music and the user's voice. This entire process can be initiated by entering a prompt such as, "Please analyze the emotions from photos and videos of our family trip, add popular music from the destination, and generate an original video with narration in my voice."

[0453] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0454] Step 1:

[0455] The server retrieves image and video data from the user's digital device or cloud storage. The input for this step is the path or URL of the data specified by the user. The server accesses the data using cloud storage APIs and downloads the files. As output, the retrieved photos and videos are stored in the server's temporary storage.

[0456] Step 2:

[0457] The server analyzes the metadata of the acquired images and videos and stores it in a database. The input is the media files acquired in step 1. The server extracts the date and time of capture and location information from these files and records it in the database. The output is organized image and video data with metadata.

[0458] Step 3:

[0459] The server uses image recognition technology to evaluate the emotions of the subject. The input for this step is the metadata-enhanced image and video data organized in step 2. The server analyzes the subject's facial expressions and atmosphere via an AI model and quantifies their emotional state. The output is a dataset with emotion tags assigned to each media file.

[0460] Step 4:

[0461] The server generates sound and narration based on emotion tags using a generative AI model. The input is a dataset to which emotion tags have been assigned in step 3. The server generates melodies that match the emotion using an application program and creates narration that mimics the user's voice using speech synthesis technology. The output is a music file and a narration file.

[0462] Step 5:

[0463] The server performs video editing, incorporating relevant background information based on the time period in which the digital data was collected. The inputs for this step are the audio and narration generated in step 4, and the metadata organized in step 2. The server evaluates and incorporates music and news that were popular during the collection period. The output is the edited original video file, with all elements integrated.

[0464] Step 6:

[0465] The device notifies the user that the edited video sent from the server is complete. The input is the edited video data, which is the output of step 5. The device uses a notification API to inform the user that the video is complete and provides a download link. The user can watch the video via this notification and relive their memories.

[0466] (Application Example 1)

[0467] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] In today's world, transforming the vast amount of photos and videos users capture daily into emotionally resonant visual experiences requires considerable time and effort. Furthermore, there is a lack of ways to enjoy this generated content in a more immersive form, not only in traditional two-dimensional formats, but also through augmented reality. Solving these challenges will enable users to consume their digital content as a richer experience.

[0469] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0470] In this invention, the server includes means for acquiring photographic and video information, means for analyzing the emotions of a subject based on the information, means for generating music and audio commentary that matches the emotions of the subject, means for incorporating the relevant environment using metadata collected during information gathering, means for editing and generating original video by combining the analysis results and generated information, and means for providing the generated video in augmented reality format. As a result, users can visually experience emotionally moving videos automatically generated based on their own photos and videos, and further gain deeper emotion and understanding through augmented reality.

[0471] "Photographic and video information" refers to still images and video data acquired using digital devices.

[0472] "Subject's emotions" refers to information that indicates a specific emotional state, obtained by analyzing the emotions of people, animals, etc., captured in photographs and videos using an AI model.

[0473] "Music and audio commentary" refers to a combination of sounds and narration that match the emotions of the subject and the related story.

[0474] "Metadata" refers to data that includes supplementary information such as the date and time a photo or video was taken, location information, and camera settings.

[0475] "Relevant environment" refers to contextual information used to incorporate the cultural, historical, and geographical background of the time of filming into the video work.

[0476] "Original video" refers to a video work that has been independently edited and generated based on the analysis results and generated content.

[0477] "Augmented reality" refers to a technological format that overlays digital content onto the real world.

[0478] The system for implementing this invention begins by transmitting photographic and video information acquired by the user's digital device to a server. The server uses this information to perform analysis and detect and classify the emotions of the subjects. An AI model is used for emotion analysis, performing calculations to identify emotion patterns. This process yields emotional data from still images and videos.

[0479] Next, the server automatically generates music and audio commentary that matches the subject based on the acquired emotional data. This process uses a generative AI model to create music segments and audio narration that correspond to the analyzed emotional state.

[0480] Furthermore, the server uses metadata associated with the information collected, such as the date and time of shooting and location information, to build a related environment. This related environment includes the cultural background and geographical elements of that period, allowing the video to reflect a sense of time and place. Based on this, a unique video for the user is generated.

[0481] The generated video is then provided to the user's device in augmented reality format. The device displays this augmented reality video, providing the user with a more immersive experience.

[0482] As a concrete example, consider a scenario where a user uploads photos of their child's sports day to a server. The server performs sentiment analysis on these photos, detecting energetic and lively emotions. Based on this, it generates upbeat music and creates voice narration suitable for parental encouragement. Furthermore, it incorporates weather information for the day of the sports day and the cultural background of the specific location as relevant environmental elements into the video. Finally, the generated video is delivered in augmented reality format through a smartphone application. This is achieved by the prompt statement, "Generate an energetic and lively video based on photos of a child's sports day, and incorporate background information specific to that day."

[0483] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0484] Step 1:

[0485] Users upload photos and videos taken with their digital devices to the server. The input consists of the user's digital media files. The server receives this data and prepares it for analysis.

[0486] Step 2:

[0487] The server extracts metadata from uploaded photos and videos. This metadata includes the date and time of capture and location information. This input metadata is used later to build the associated environment. The server records this information in a database.

[0488] Step 3:

[0489] The server uses a generative AI model to analyze the emotions of subjects in photos and videos. Image data is used as input, and the output is estimated emotion information of the subjects. The generative AI model used in this process is trained using TensorFlow or PyTorch.

[0490] Step 4:

[0491] The server generates music and audio commentary that match the estimated emotions based on the results of emotion analysis. The input is analyzed emotion information, and the generating AI model outputs music and audio narration. The audio narration is generated by the AI ​​mimicking the user's voice.

[0492] Step 5:

[0493] The server constructs a related environment based on the extracted metadata, corresponding to the date and time of shooting and location information. This related environment includes cultural or geographical background information and provides information for incorporation into the video. Shooting information is used as input, and related environment data is generated as output.

[0494] Step 6:

[0495] The server combines the acquired emotion analysis results, generated music and audio commentary, and relevant environmental data to edit and produce original video. All intermediate products are inputs, and the final video is provided as output.

[0496] Step 7:

[0497] The device notifies the user of the final video content sent from the server. The user receives this video, which is provided in augmented reality format, and by experiencing the video displayed on the device, they can enjoy a more immersive experience of their memories.

[0498] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0499] This invention is a system that uses an emotion engine to recognize the user's emotions based on user-saved photo and video data, and highly personalizes the generation of video content. This system functions primarily through a series of processes performed on a server and user interaction.

[0500] First, the server securely retrieves photo and video data from the user's digital storage and aggregates metadata, including the date and time of capture and location information. This data is temporarily stored on the server in preparation for the next processing step.

[0501] Next, the server uses an AI model to analyze the emotions of the subjects in the acquired photos and videos. In addition, it utilizes an emotion engine to collect the user's voice, facial expressions, gaze, etc., in real time through sensors, and incorporates the user's own emotional state into the system.

[0502] Furthermore, the server generates music and narration appropriate to each scene of the video based on emotion analysis and data derived from user emotions. In this process, AI technology is used to mimic the user's voice in the narration, resulting in a more relatable narrative. Additionally, based on metadata collected during data collection, the system incorporates context into the video, considering the historical background and trends of the time of filming.

[0503] Here, the server aggregates this information, drives the editing software, and generates the final original video. It can also adjust the tempo and atmosphere of the video as needed, reflecting the results of real-time sentiment analysis of the user.

[0504] Ultimately, the device notifies the user of the generated video, and the user views the video on their device. This video incorporates the user's emotions and historical context, providing a new way to revisit family memories. For example, when the user expresses joy, the mood of the video is brightened, and the narration is adjusted to reflect that emotion. In this way, the present invention provides advanced personalization that takes the user's emotions into consideration, supporting the creation of useful video content that deepens family bonds.

[0505] The following describes the processing flow.

[0506] Step 1:

[0507] The server periodically collects photos and videos stored by users in their digital storage. During this process, metadata such as the date and time of capture and location information is also acquired. The data is stored in temporary storage for analysis, while maintaining security considerations.

[0508] Step 2:

[0509] The server analyzes the acquired photo and video data using an emotion engine. This process calculates an emotion score from the facial expressions and actions of subjects in the images and videos, and further infers and records the user's emotions using voice input from the user and audio, facial, and gaze data collected in real time through the device's camera.

[0510] Step 3:

[0511] The server generates music and AI narration appropriate for the video scene based on the analyzed emotion data. The narration is generated using AI technology that has learned from the user's past voice data, mimicking the user's voice. The music is created by adjusting the tempo and key to match the analyzed emotions.

[0512] Step 4:

[0513] The server uses aggregated metadata and user sentiment information to reflect historical data and trends from the time of filming in the video. This information influences the selection of narration and music within the video and is used to provide overall context.

[0514] Step 5:

[0515] The server uses editing software to edit photo and video data in conjunction with sentiment analysis and generated content. Generated music and narration are appropriately inserted into video scenes to create emotionally rich videos. Here, the tempo and mood of the video are dynamically adjusted based on the user's real-time sentiment data.

[0516] Step 6:

[0517] The device notifies the user when the generated video is ready. The user can watch the video through the application and choose between streaming or downloading. During viewing, the user can experience the inclusion of historical data based on past events and emotionally nuanced elements.

[0518] (Example 2)

[0519] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0520] With the widespread use of digital photos and videos, there is a challenge in how to individually optimize the vast amount of media data accumulated by users and provide a deeper emotional experience. Furthermore, existing systems have a problem in that they do not adequately generate personalized content that reflects the user's emotional state in real time.

[0521] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0522] In this invention, the server includes a device for acquiring photographic and video data, a device for analyzing the emotions of a subject based on the data, and a device for generating music and narration based on the emotions of the subject and the emotional state of the user. This makes it possible to generate highly personalized video content that matches the user's emotions.

[0523] A "device for acquiring photo and video data" is a device for securely transferring photos and videos from a user's digital storage.

[0524] A "device for analyzing the emotions of subjects" is a device that uses an AI model to detect and analyze the emotions of subjects appearing in photographs and videos.

[0525] A "device that generates music and narration based on the user's emotional state" is a device that generates personalized music and narration based on user emotional data acquired in real time.

[0526] A "device that incorporates related information by utilizing metadata acquired during data acquisition" is a device that analyzes the metadata of photos and videos (for example, the date and time of shooting and location information) and incorporates the associated related information into the video content.

[0527] A "device for integrating analysis results and generated content to edit and generate personalized videos" is a device that uses the results of sentiment analysis to edit content, including music and narration, into a video, providing a personalized viewing experience for each user.

[0528] A "device that adjusts based on real-time emotional state" is a device that adjusts the tempo and atmosphere of a video by reflecting the user's current emotions during video generation.

[0529] A "device that provides and makes accessible to users" is a device that notifies users of the generated video and allows them to easily access and watch it on their own devices.

[0530] This invention is a system that generates highly personalized video content based on photos and videos owned by the user, while taking into account their emotional state. The system functions primarily through interaction between a server, a terminal, and the user. The respective components and specific embodiments are described below.

[0531] The server is the primary device for retrieving photo and video data from the user's digital storage. Network communication protocols are used for secure data retrieval. The retrieved data, along with metadata (e.g., date and location information), is temporarily stored on the server. Based on this data, an AI model is used to analyze the emotions of the subjects. The AI ​​model employs image recognition technology. Specifically, it identifies emotions such as smiles and sadness based on the subject's facial expressions.

[0532] Furthermore, the server collects the user's voice, facial expressions, and gaze in real time through sensors (camera, microphone, etc.) installed in the user's device. This activates the emotion engine, integrating the user's real-time emotional state into the system. Based on this information, AI technology is used to generate music and narration that mimics the user's voice. The generated music and narration are designed to harmonize with the user's emotions at that moment.

[0533] When generating narration, a prompt like the following might be used: "Based on the sentiment analysis results of the photos and videos, generate narration that reflects the user's feelings of joy. The narration should match the user's voice tone and have a cheerful mood."

[0534] Furthermore, the server integrates these analysis results and content, and uses editing software to generate the final video. During this process, historical data, built based on shooting date and location information, is used to enhance the scenes, providing richer context to the video.

[0535] Ultimately, the device notifies the user of the generated video, which the user can then freely view on their device. The video, which reflects the user's emotions and surrounding environment, offers a new experience for revisiting memories with family and loved ones.

[0536] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0537] Step 1:

[0538] The server obtains permission from the user to connect to digital storage and retrieves photo and video data. The input is the user's storage path or cloud storage access information. The retrieved data is stored on the server along with metadata such as the date and time of capture and location information. During this process, encryption protocols such as SSL are used to ensure security during data transfer.

[0539] Step 2:

[0540] The server inputs collected photo and video data into an AI model to analyze the emotions of the subjects. The input is image and video data, and the output is data indicating emotional characteristics such as smiles, sadness, and surprise. Specifically, it uses image recognition technology to identify facial features and machine learning algorithms to infer specific emotions.

[0541] Step 3:

[0542] The server collects voice, gaze, and facial expressions in real time from sensors via the user's device. The input is raw data from the microphone and camera, which is used to infer the user's current emotional state. The output is emotional data, such as whether the user is happy or relaxed. An emotion engine is used to process this information, and the analysis results are instantly reflected on the server.

[0543] Step 4:

[0544] The server generates music and narration based on the results of sentiment analysis and the user's emotional state. The input consists of sentiment data and prompt text perceived by the program. The output is music and narration with a tone and content that matches the user's emotions. Using a generative AI model, the narration is created in a natural tone based on the user's past voice data.

[0545] Step 5:

[0546] The server integrates analysis results and generated content, and uses editing software to produce personalized videos. Inputs include music, narration, and original image / video data, while output is the completed original video. This process incorporates historical context and trends based on metadata from the time of filming, striving for visual aesthetics appropriate to the content.

[0547] Step 6:

[0548] The device notifies the user of the generated video and provides it in a playable format. The input is video data from the server, and the output is video content played on the device's screen. Users can easily play this video on their own devices and enjoy a personalized viewing experience.

[0549] (Application Example 2)

[0550] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0551] When generating personalized video content from digital data, it is necessary to go beyond simply relying on static sentiment analysis and instead reflect the user's real-time emotional state to provide more relatable and empathetic content. Furthermore, it is essential to provide more personal and approachable narration to enhance the memorable experience.

[0552] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0553] In this invention, the server includes means for acquiring digital information, means for analyzing the emotions of a target based on the information, and means for collecting the user's emotional state in real time and adjusting the tempo and atmosphere in the video. This makes it possible to generate video that more faithfully responds to the user's emotions.

[0554] "Digital information" refers to visual media data stored electronically, such as photographs and videos.

[0555] "Subject" is a term that refers to a person or object that appears in a photograph or video.

[0556] "Analyzing emotions" means estimating a person's emotional state from their facial expressions and gestures in visual media.

[0557] "Musical material and explanatory text" refers to background music suitable for the video and text generated as narration, which are added to the visual content.

[0558] "Auxiliary information" refers to metadata such as location information and date / time information that is collected when taking photos or videos.

[0559] "Background" refers to the historical context and cultural background that are incorporated to make the video content easier to understand.

[0560] "Unique video content" refers to personalized video content generated based on the user's specific emotions.

[0561] "User" is a term that refers to an individual or organization using a digital information provision system.

[0562] "Real-time data collection" refers to the process by which emotional data is instantly acquired as the user uses the system.

[0563] "Adjusting tempo and atmosphere" means dynamically changing the playback speed of the video and the tone of the music to match the user's emotions.

[0564] To realize this invention, the server first acquires photo and video data from the user's digital device. The data includes metadata such as the date and time of shooting and location information, and this is securely stored as digital information in a cloud environment. The server performs sentiment analysis on the image and video data using an AI model. In this process, the subject's facial expressions and behavioral patterns are analyzed, and their emotional state is estimated. Furthermore, the tempo and atmosphere of the video are dynamically adjusted based on the user's sentiment data collected in real time.

[0565] On the software side, the server uses sentiment analysis engines such as Google Cloud AI and Amazon Rekognition to process data. These engines extract emotional information from fragmented data and derive the most appropriate personalization elements. Based on the analysis results, the server generates music and explanatory text, and a generative AI model provides user-specific narration. The narration enhances personal familiarity by mimicking the user's past voice data.

[0566] As a concrete example, consider a case where photos from a family trip are uploaded to a server. In this case, if the analyzed emotion is determined to be "joy," the server will generate music that creates a joyful atmosphere, along with narration reflecting on the feelings at the time of the trip. Once the video is complete, the device will notify the user of the generated video, making it available for viewing.

[0567] An example of a prompt for a generative AI model is, "Extract fun moments from a family photo album and create an exciting music video."

[0568] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0569] Step 1:

[0570] Users upload photos and videos to the cloud using their smartphones. The input consists of photos and videos taken from the user's digital device. The output is that this data is stored securely on a server.

[0571] Step 2:

[0572] The server analyzes metadata (such as date and location information) of photos and videos stored in the cloud. The input is the metadata of photos and videos uploaded by the user. The output is the analyzed metadata. This data is later used to build context for related background information.

[0573] Step 3:

[0574] The server uses an AI model to analyze the emotions of subjects from photos and videos. The input is visual media data from the user. The AI ​​model analyzes facial expressions and behavioral patterns, and generates data that identifies the emotions of the subjects as output.

[0575] Step 4:

[0576] The server collects user voice and facial expression data in real time and adjusts the video tempo and atmosphere based on that data. The input is the user's real-time voice and facial expression data. The output is adjusted video settings that take the user's emotions into consideration.

[0577] Step 5:

[0578] The server generates music and explanatory text based on the analysis results. The input is the subject's emotional data, and the output is music and narration text appropriate to that emotion.

[0579] Step 6:

[0580] The server uses a generated AI model to provide user-specific narration. Inputs include the user's past voice data and narration text. Output is narration audio that closely resembles the user's voice.

[0581] Step 7:

[0582] The server combines all elements (adjusted video, music, narration) to edit and generate a unique video. Inputs are the adjusted video settings, music, and narration. The output is a fully edited, personalized video.

[0583] Step 8:

[0584] The device notifies the user of the completed video and makes it available for viewing. The input is the edited video data. As output, the user can view the video generated on their device.

[0585] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0586] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0587] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0588] [Fourth Embodiment]

[0589] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0590] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0591] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0592] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0593] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0594] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0595] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0596] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0597] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0598] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0599] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0600] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0601] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0602] This invention is a system that utilizes photographic and video data saved by users using digital devices to automatically generate emotionally impactful video works. This system consists mainly of a series of processes performed by a server.

[0603] First, the server organizes the photo and video data retrieved from the user's digital storage. This includes verifying and recording metadata such as the date and time of shooting and location information.

[0604] Next, the server uses an AI model to analyze the emotions of subjects in photos and videos based on the acquired data. This analysis is then used to provide a means for generating music appropriate for each scene in the video.

[0605] Subsequently, the server automatically generates music and narration tailored to the user's personality based on the information obtained from the emotion analysis. For the narration, an AI that has learned from the user's existing data is able to mimic the user's voice to tell the story. This allows family memories to be expressed from a more individual perspective.

[0606] Furthermore, the server incorporates relevant context based on the date and location information of the footage, for example, by including music or news that was popular at the time, providing a rich narrative that reflects the historical context.

[0607] Ultimately, the server edits the generated music, narration, and context to create an original video. This video can offer users a new perspective and visually reconstruct the family story.

[0608] The device notifies the user of edited videos sent from the server, and the user can watch and download them on their own device to relive memories with family. For example, if the video captures a family trip, the emotions captured in the photos and videos taken during the trip will be analyzed and presented as a travelogue narrated by the user, along with music that was popular at the travel destination.

[0609] The above is an example of an embodiment for carrying out the present invention. This system allows users to easily enjoy high-quality family history from their saved data.

[0610] The following describes the processing flow.

[0611] Step 1:

[0612] The server periodically retrieves photo and video data from the user's device or cloud storage. Along with this data, it also collects metadata such as the date and time of capture and location information. The collected data is stored in temporary storage and prepared for analysis.

[0613] Step 2:

[0614] The server inputs the acquired photo and video data into an AI model to analyze the emotions of the subjects. This model analyzes facial expressions and actions within images and videos to generate emotion scores (e.g., joy, sadness, surprise) for each image or video. The analysis results are recorded in a database and used during video editing.

[0615] Step 3:

[0616] The server generates music appropriate for each scene in the video based on the results of emotion analysis. In this music generation process, the tempo and mood of the music are adjusted according to the emotion score, based on pre-prepared music templates. At the same time, an AI narrator learns from the user's past voice data and creates narration that mimics the user's voice.

[0617] Step 4:

[0618] The server utilizes metadata collected during data collection to investigate the historical context, such as popular music and news at the time of filming. Based on this information, it prepares materials to give the video a historical background and narrative. This background context information is used during the final video editing process.

[0619] Step 5:

[0620] The server uses editing software to combine the acquired photos and videos with analysis results and generated content to create the original video. Each media file is placed on a timeline according to the date and time of shooting, and emotionally appropriate background music and narration are inserted. This creates a consistent storyline.

[0621] Step 6:

[0622] The device notifies the user that the completed video is ready in their account. The user can preview the video through the application and play it by selecting streaming or download options. The user then has the opportunity to watch the generated video with family and reminisce about memories.

[0623] (Example 1)

[0624] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0625] In modern times, many people generate vast amounts of photographic and video data using digital devices. However, meaningfully utilizing this data and reconstructing it into emotionally impactful visual works is a significant challenge. Furthermore, there is a need for methods that can effectively express the emotions embedded in individual data and the cultural context of the time period.

[0626] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0627] In this invention, the server includes a mechanism for acquiring image and video information, means for evaluating the emotions of the subject, and technology for generating sounds and narration that are appropriate to the evaluation results. This makes it possible to automatically generate emotionally rich video works from the user's digital data and to visually and emotionally reconstruct individual memories.

[0628] "Image and video information" refers to visual media data such as photographs and videos acquired by digital devices.

[0629] "Evaluating the emotions of a subject" refers to the process of analyzing the facial expressions and atmosphere of a subject in a photograph or video and quantifying their emotional state.

[0630] "Technology for generating sound and narration" refers to technology that automatically creates suitable music and narration based on the results of an emotional evaluation of a subject.

[0631] A "mechanism" refers to a system or device designed to perform a specific function.

[0632] "Auxiliary information" refers to metadata such as date, time, and location information that accompanies data collection.

[0633] "Relevant background information" refers to historical and cultural data used to supplement the context of captured images and videos.

[0634] "Editing and creation methods" refers to the entire process of combining acquired data with generated content to construct a new video work.

[0635] "Original video" refers to video works that are individually produced to express a specific user experience or story.

[0636] This invention is a system that utilizes image and video information saved by users using digital devices and automatically generates emotionally impactful video works based on them. The system is primarily processed on a server and consists of the following hardware and software.

[0637] The server retrieves image and video data from the user's digital device or cloud storage. To achieve this, it utilizes cloud storage APIs and has the capability to access data and download necessary files. Next, the server organizes the retrieved data along with metadata for the photos and videos. A database system is used for analyzing the metadata, recording the date and time of capture and location information.

[0638] The server uses an AI model to analyze the emotions of the subject. It incorporates a process that utilizes image recognition technology to identify the subject and quantify its emotional state. The results of the emotional analysis are tagged and used to generate music and narration.

[0639] The server then uses a generative AI model to create sounds tailored to the user's personality. The music generation employs algorithms to generate melodies that match emotions, and by mimicking the user's voice using speech synthesis technology, it enables narration with a more personalized perspective.

[0640] Furthermore, the server incorporates relevant background information based on the date and time of shooting and location data, for example, by reflecting music and news that were popular at the time the data was collected, thereby providing a richer narrative. This video editing blends multiple elements of music, narration, and emotionally rich background information.

[0641] The device receives the edited video sent from the server and notifies the user. The user can then watch or download the video via this notification and relive personal memories.

[0642] As a concrete example, in a video summarizing a family trip, the system analyzes the emotions captured in photos and videos taken during the trip and generates a travelogue narrated with appropriate background music and the user's voice. This entire process can be initiated by entering a prompt such as, "Please analyze the emotions from photos and videos of our family trip, add popular music from the destination, and generate an original video with narration in my voice."

[0643] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0644] Step 1:

[0645] The server retrieves image and video data from the user's digital device or cloud storage. The input for this step is the path or URL of the data specified by the user. The server accesses the data using cloud storage APIs and downloads the files. As output, the retrieved photos and videos are stored in the server's temporary storage.

[0646] Step 2:

[0647] The server analyzes the metadata of the acquired images and videos and stores it in a database. The input is the media files acquired in step 1. The server extracts the date and time of capture and location information from these files and records it in the database. The output is organized image and video data with metadata.

[0648] Step 3:

[0649] The server uses image recognition technology to evaluate the emotions of the subject. The input for this step is the metadata-enhanced image and video data organized in step 2. The server analyzes the subject's facial expressions and atmosphere via an AI model and quantifies their emotional state. The output is a dataset with emotion tags assigned to each media file.

[0650] Step 4:

[0651] The server generates sound and narration based on emotion tags using a generative AI model. The input is a dataset to which emotion tags have been assigned in step 3. The server generates melodies that match the emotion using an application program and creates narration that mimics the user's voice using speech synthesis technology. The output is a music file and a narration file.

[0652] Step 5:

[0653] The server performs video editing, incorporating relevant background information based on the time period in which the digital data was collected. The inputs for this step are the audio and narration generated in step 4, and the metadata organized in step 2. The server evaluates and incorporates music and news that were popular during the collection period. The output is the edited original video file, with all elements integrated.

[0654] Step 6:

[0655] The device notifies the user that the edited video sent from the server is complete. The input is the edited video data, which is the output of step 5. The device uses a notification API to inform the user that the video is complete and provides a download link. The user can watch the video via this notification and relive their memories.

[0656] (Application Example 1)

[0657] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0658] In today's world, transforming the vast amount of photos and videos users capture daily into emotionally resonant visual experiences requires considerable time and effort. Furthermore, there is a lack of ways to enjoy this generated content in a more immersive form, not only in traditional two-dimensional formats, but also through augmented reality. Solving these challenges will enable users to consume their digital content as a richer experience.

[0659] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0660] In this invention, the server includes means for acquiring photographic and video information, means for analyzing the emotions of a subject based on the information, means for generating music and audio commentary that matches the emotions of the subject, means for incorporating the relevant environment using metadata collected during information gathering, means for editing and generating original video by combining the analysis results and generated information, and means for providing the generated video in augmented reality format. As a result, users can visually experience emotionally moving videos automatically generated based on their own photos and videos, and further gain deeper emotion and understanding through augmented reality.

[0661] "Photographic and video information" refers to still images and video data acquired using digital devices.

[0662] "Subject's emotions" refers to information that indicates a specific emotional state, obtained by analyzing the emotions of people, animals, etc., captured in photographs and videos using an AI model.

[0663] "Music and audio commentary" refers to a combination of sounds and narration that match the emotions of the subject and the related story.

[0664] "Metadata" refers to data that includes supplementary information such as the date and time a photo or video was taken, location information, and camera settings.

[0665] "Relevant environment" refers to contextual information used to incorporate the cultural, historical, and geographical background of the time of filming into the video work.

[0666] "Original video" refers to a video work that has been independently edited and generated based on the analysis results and generated content.

[0667] "Augmented reality" refers to a technological format that overlays digital content onto the real world.

[0668] The system for implementing this invention begins by transmitting photographic and video information acquired by the user's digital device to a server. The server uses this information to perform analysis and detect and classify the emotions of the subjects. An AI model is used for emotion analysis, performing calculations to identify emotion patterns. This process yields emotional data from still images and videos.

[0669] Next, the server automatically generates music and audio commentary that matches the subject based on the acquired emotional data. This process uses a generative AI model to create music segments and audio narration that correspond to the analyzed emotional state.

[0670] Furthermore, the server uses metadata associated with the information collected, such as the date and time of shooting and location information, to build a related environment. This related environment includes the cultural background and geographical elements of that period, allowing the video to reflect a sense of time and place. Based on this, a unique video for the user is generated.

[0671] The generated video is then provided to the user's device in augmented reality format. The device displays this augmented reality video, providing the user with a more immersive experience.

[0672] As a concrete example, consider a scenario where a user uploads photos of their child's sports day to a server. The server performs sentiment analysis on these photos, detecting energetic and lively emotions. Based on this, it generates upbeat music and creates voice narration suitable for parental encouragement. Furthermore, it incorporates weather information for the day of the sports day and the cultural background of the specific location as relevant environmental elements into the video. Finally, the generated video is delivered in augmented reality format through a smartphone application. This is achieved by the prompt statement, "Generate an energetic and lively video based on photos of a child's sports day, and incorporate background information specific to that day."

[0673] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0674] Step 1:

[0675] Users upload photos and videos taken with their digital devices to the server. The input consists of the user's digital media files. The server receives this data and prepares it for analysis.

[0676] Step 2:

[0677] The server extracts metadata from uploaded photos and videos. This metadata includes the date and time of capture and location information. This input metadata is used later to build the associated environment. The server records this information in a database.

[0678] Step 3:

[0679] The server uses a generative AI model to analyze the emotions of subjects in photos and videos. Image data is used as input, and the output is estimated emotion information of the subjects. The generative AI model used in this process is trained using TensorFlow or PyTorch.

[0680] Step 4:

[0681] The server generates music and audio commentary that match the estimated emotions based on the results of emotion analysis. The input is analyzed emotion information, and the generating AI model outputs music and audio narration. The audio narration is generated by the AI ​​mimicking the user's voice.

[0682] Step 5:

[0683] The server constructs a related environment based on the extracted metadata, corresponding to the date and time of shooting and location information. This related environment includes cultural or geographical background information and provides information for incorporation into the video. Shooting information is used as input, and related environment data is generated as output.

[0684] Step 6:

[0685] The server combines the acquired emotion analysis results, generated music and audio commentary, and relevant environmental data to edit and produce original video. All intermediate products are inputs, and the final video is provided as output.

[0686] Step 7:

[0687] The device notifies the user of the final video content sent from the server. The user receives this video, which is provided in augmented reality format, and by experiencing the video displayed on the device, they can enjoy a more immersive experience of their memories.

[0688] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0689] This invention is a system that uses an emotion engine to recognize the user's emotions based on user-saved photo and video data, and highly personalizes the generation of video content. This system functions primarily through a series of processes performed on a server and user interaction.

[0690] First, the server securely retrieves photo and video data from the user's digital storage and aggregates metadata, including the date and time of capture and location information. This data is temporarily stored on the server in preparation for the next processing step.

[0691] Next, the server uses an AI model to analyze the emotions of the subjects in the acquired photos and videos. In addition, it utilizes an emotion engine to collect the user's voice, facial expressions, gaze, etc., in real time through sensors, and incorporates the user's own emotional state into the system.

[0692] Furthermore, the server generates music and narration appropriate to each scene of the video based on emotion analysis and data derived from user emotions. In this process, AI technology is used to mimic the user's voice in the narration, resulting in a more relatable narrative. Additionally, based on metadata collected during data collection, the system incorporates context into the video, considering the historical background and trends of the time of filming.

[0693] Here, the server aggregates this information, drives the editing software, and generates the final original video. It can also adjust the tempo and atmosphere of the video as needed, reflecting the results of real-time sentiment analysis of the user.

[0694] Ultimately, the device notifies the user of the generated video, and the user views the video on their device. This video incorporates the user's emotions and historical context, providing a new way to revisit family memories. For example, when the user expresses joy, the mood of the video is brightened, and the narration is adjusted to reflect that emotion. In this way, the present invention provides advanced personalization that takes the user's emotions into consideration, supporting the creation of useful video content that deepens family bonds.

[0695] The following describes the processing flow.

[0696] Step 1:

[0697] The server periodically collects photos and videos stored by users in their digital storage. During this process, metadata such as the date and time of capture and location information is also acquired. The data is stored in temporary storage for analysis, while maintaining security considerations.

[0698] Step 2:

[0699] The server analyzes the acquired photo and video data using an emotion engine. This process calculates an emotion score from the facial expressions and actions of subjects in the images and videos, and further infers and records the user's emotions using voice input from the user and audio, facial, and gaze data collected in real time through the device's camera.

[0700] Step 3:

[0701] The server generates music and AI narration appropriate for the video scene based on the analyzed emotion data. The narration is generated using AI technology that has learned from the user's past voice data, mimicking the user's voice. The music is created by adjusting the tempo and key to match the analyzed emotions.

[0702] Step 4:

[0703] The server uses aggregated metadata and user sentiment information to reflect historical data and trends from the time of filming in the video. This information influences the selection of narration and music within the video and is used to provide overall context.

[0704] Step 5:

[0705] The server uses editing software to edit photo and video data in conjunction with sentiment analysis and generated content. Generated music and narration are appropriately inserted into video scenes to create emotionally rich videos. Here, the tempo and mood of the video are dynamically adjusted based on the user's real-time sentiment data.

[0706] Step 6:

[0707] The device notifies the user when the generated video is ready. The user can watch the video through the application and choose between streaming or downloading. During viewing, the user can experience the inclusion of historical data based on past events and emotionally nuanced elements.

[0708] (Example 2)

[0709] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0710] With the widespread use of digital photos and videos, there is a challenge in how to individually optimize the vast amount of media data accumulated by users and provide a deeper emotional experience. Furthermore, existing systems have a problem in that they do not adequately generate personalized content that reflects the user's emotional state in real time.

[0711] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0712] In this invention, the server includes a device for acquiring photographic and video data, a device for analyzing the emotions of a subject based on the data, and a device for generating music and narration based on the emotions of the subject and the emotional state of the user. This makes it possible to generate highly personalized video content that matches the user's emotions.

[0713] A "device for acquiring photo and video data" is a device for securely transferring photos and videos from a user's digital storage.

[0714] A "device for analyzing the emotions of subjects" is a device that uses an AI model to detect and analyze the emotions of subjects appearing in photographs and videos.

[0715] A "device that generates music and narration based on the user's emotional state" is a device that generates personalized music and narration based on user emotional data acquired in real time.

[0716] A "device that incorporates related information by utilizing metadata acquired during data acquisition" is a device that analyzes the metadata of photos and videos (for example, the date and time of shooting and location information) and incorporates the associated related information into the video content.

[0717] A "device for integrating analysis results and generated content to edit and generate personalized videos" is a device that uses the results of sentiment analysis to edit content, including music and narration, into a video, providing a personalized viewing experience for each user.

[0718] A "device that adjusts based on real-time emotional state" is a device that adjusts the tempo and atmosphere of a video by reflecting the user's current emotions during video generation.

[0719] A "device that provides and makes accessible to users" is a device that notifies users of the generated video and allows them to easily access and watch it on their own devices.

[0720] This invention is a system that generates highly personalized video content based on photos and videos owned by the user, while taking into account their emotional state. The system functions primarily through interaction between a server, a terminal, and the user. The respective components and specific embodiments are described below.

[0721] The server is the primary device for retrieving photo and video data from the user's digital storage. Network communication protocols are used for secure data retrieval. The retrieved data, along with metadata (e.g., date and location information), is temporarily stored on the server. Based on this data, an AI model is used to analyze the emotions of the subjects. The AI ​​model employs image recognition technology. Specifically, it identifies emotions such as smiles and sadness based on the subject's facial expressions.

[0722] Furthermore, the server collects the user's voice, facial expressions, and gaze in real time through sensors (camera, microphone, etc.) installed in the user's device. This activates the emotion engine, integrating the user's real-time emotional state into the system. Based on this information, AI technology is used to generate music and narration that mimics the user's voice. The generated music and narration are designed to harmonize with the user's emotions at that moment.

[0723] When generating narration, a prompt like the following might be used: "Based on the sentiment analysis results of the photos and videos, generate narration that reflects the user's feelings of joy. The narration should match the user's voice tone and have a cheerful mood."

[0724] Furthermore, the server integrates these analysis results and content, and uses editing software to generate the final video. During this process, historical data, built based on shooting date and location information, is used to enhance the scenes, providing richer context to the video.

[0725] Ultimately, the device notifies the user of the generated video, which the user can then freely view on their device. The video, which reflects the user's emotions and surrounding environment, offers a new experience for revisiting memories with family and loved ones.

[0726] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0727] Step 1:

[0728] The server obtains permission from the user to connect to digital storage and retrieves photo and video data. The input is the user's storage path or cloud storage access information. The retrieved data is stored on the server along with metadata such as the date and time of capture and location information. During this process, encryption protocols such as SSL are used to ensure security during data transfer.

[0729] Step 2:

[0730] The server inputs collected photo and video data into an AI model to analyze the emotions of the subjects. The input is image and video data, and the output is data indicating emotional characteristics such as smiles, sadness, and surprise. Specifically, it uses image recognition technology to identify facial features and machine learning algorithms to infer specific emotions.

[0731] Step 3:

[0732] The server collects voice, gaze, and facial expressions in real time from sensors via the user's device. The input is raw data from the microphone and camera, which is used to infer the user's current emotional state. The output is emotional data, such as whether the user is happy or relaxed. An emotion engine is used to process this information, and the analysis results are instantly reflected on the server.

[0733] Step 4:

[0734] The server generates music and narration based on the results of sentiment analysis and the user's emotional state. The input consists of sentiment data and prompt text perceived by the program. The output is music and narration with a tone and content that matches the user's emotions. Using a generative AI model, the narration is created in a natural tone based on the user's past voice data.

[0735] Step 5:

[0736] The server integrates analysis results and generated content, and uses editing software to produce personalized videos. Inputs include music, narration, and original image / video data, while output is the completed original video. This process incorporates historical context and trends based on metadata from the time of filming, striving for visual aesthetics appropriate to the content.

[0737] Step 6:

[0738] The device notifies the user of the generated video and provides it in a playable format. The input is video data from the server, and the output is video content played on the device's screen. Users can easily play this video on their own devices and enjoy a personalized viewing experience.

[0739] (Application Example 2)

[0740] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0741] When generating personalized video content from digital data, it is necessary to go beyond simply relying on static sentiment analysis and instead reflect the user's real-time emotional state to provide more relatable and empathetic content. Furthermore, it is essential to provide more personal and approachable narration to enhance the memorable experience.

[0742] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0743] In this invention, the server includes means for acquiring digital information, means for analyzing the emotions of a target based on the information, and means for collecting the user's emotional state in real time and adjusting the tempo and atmosphere in the video. This makes it possible to generate video that more faithfully responds to the user's emotions.

[0744] "Digital information" refers to visual media data stored electronically, such as photographs and videos.

[0745] "Subject" is a term that refers to a person or object that appears in a photograph or video.

[0746] "Analyzing emotions" means estimating a person's emotional state from their facial expressions and gestures in visual media.

[0747] "Musical material and explanatory text" refers to background music suitable for the video and text generated as narration, which are added to the visual content.

[0748] "Auxiliary information" refers to metadata such as location information and date / time information that is collected when taking photos or videos.

[0749] "Background" refers to the historical context and cultural background that are incorporated to make the video content easier to understand.

[0750] "Unique video content" refers to personalized video content generated based on the user's specific emotions.

[0751] "User" is a term that refers to an individual or organization using a digital information provision system.

[0752] "Real-time data collection" refers to the process by which emotional data is instantly acquired as the user uses the system.

[0753] "Adjusting tempo and atmosphere" means dynamically changing the playback speed of the video and the tone of the music to match the user's emotions.

[0754] To realize this invention, the server first acquires photo and video data from the user's digital device. The data includes metadata such as the date and time of shooting and location information, and this is securely stored as digital information in a cloud environment. The server performs sentiment analysis on the image and video data using an AI model. In this process, the subject's facial expressions and behavioral patterns are analyzed, and their emotional state is estimated. Furthermore, the tempo and atmosphere of the video are dynamically adjusted based on the user's sentiment data collected in real time.

[0755] On the software side, the server uses sentiment analysis engines such as Google Cloud AI and Amazon Rekognition to process data. These engines extract emotional information from fragmented data and derive the most appropriate personalization elements. Based on the analysis results, the server generates music and explanatory text, and a generative AI model provides user-specific narration. The narration enhances personal familiarity by mimicking the user's past voice data.

[0756] As a concrete example, consider a case where photos from a family trip are uploaded to a server. In this case, if the analyzed emotion is determined to be "joy," the server will generate music that creates a joyful atmosphere, along with narration reflecting on the feelings at the time of the trip. Once the video is complete, the device will notify the user of the generated video, making it available for viewing.

[0757] An example of a prompt for a generative AI model is, "Extract fun moments from a family photo album and create an exciting music video."

[0758] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0759] Step 1:

[0760] Users upload photos and videos to the cloud using their smartphones. The input consists of photos and videos taken from the user's digital device. The output is that this data is stored securely on a server.

[0761] Step 2:

[0762] The server analyzes metadata (such as date and location information) of photos and videos stored in the cloud. The input is the metadata of photos and videos uploaded by the user. The output is the analyzed metadata. This data is later used to build context for related background information.

[0763] Step 3:

[0764] The server uses an AI model to analyze the emotions of subjects from photos and videos. The input is visual media data from the user. The AI ​​model analyzes facial expressions and behavioral patterns, and generates data that identifies the emotions of the subjects as output.

[0765] Step 4:

[0766] The server collects user voice and facial expression data in real time and adjusts the video tempo and atmosphere based on that data. The input is the user's real-time voice and facial expression data. The output is adjusted video settings that take the user's emotions into consideration.

[0767] Step 5:

[0768] The server generates music and explanatory text based on the analysis results. The input is the subject's emotional data, and the output is music and narration text appropriate to that emotion.

[0769] Step 6:

[0770] The server uses a generated AI model to provide user-specific narration. Inputs include the user's past voice data and narration text. Output is narration audio that closely resembles the user's voice.

[0771] Step 7:

[0772] The server combines all elements (adjusted video, music, narration) to edit and generate a unique video. Inputs are the adjusted video settings, music, and narration. The output is a fully edited, personalized video.

[0773] Step 8:

[0774] The device notifies the user of the completed video and makes it available for viewing. The input is the edited video data. As output, the user can view the video generated on their device.

[0775] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0776] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0777] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0778] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0779] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0780] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0781] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0782] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0783] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0784] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0785] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0786] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0787] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0788] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0789] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0790] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0791] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0792] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0793] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0794] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0795] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0796] The following is further disclosed regarding the embodiments described above.

[0797] (Claim 1)

[0798] Means for acquiring photographic and video data,

[0799] A means for analyzing the emotions of the subject based on the aforementioned data,

[0800] Means for generating music and narration that match the emotions of the subject,

[0801] A means of incorporating relevant context by utilizing metadata collected during data acquisition,

[0802] A means for editing and generating an original video by combining the aforementioned analysis results and generated content,

[0803] A means of providing the aforementioned video to the user,

[0804] A system that includes this.

[0805] (Claim 2)

[0806] The system according to claim 1, wherein the narration is generated using technology that imitates the user's voice.

[0807] (Claim 3)

[0808] The system according to claim 1, wherein the aforementioned related context is constructed using historical data based on the date and time of shooting and location information.

[0809] "Example 1"

[0810] (Claim 1)

[0811] A mechanism for acquiring image and video information,

[0812] A means for evaluating the subject's emotions based on the aforementioned information,

[0813] A technology for generating sounds and narration that conform to the aforementioned evaluation results,

[0814] A means of integrating relevant background information by utilizing auxiliary information during information gathering,

[0815] A method for editing and creating original video by combining the aforementioned evaluation results and generated content,

[0816] A method for providing the aforementioned video to users,

[0817] A system that includes this.

[0818] (Claim 2)

[0819] The system according to claim 1, wherein the aforementioned narration is generated by employing a technology that reproduces the user's way of speaking.

[0820] (Claim 3)

[0821] The system according to claim 1, wherein the aforementioned related background information is formed by referring to cultural information based on the date and time and location information of the photograph.

[0822] "Application Example 1"

[0823] (Claim 1)

[0824] Means for acquiring photographic and video information,

[0825] A means for analyzing the emotions of the subject based on the aforementioned information,

[0826] A means for generating music and audio commentary that matches the emotions of the subject,

[0827] A method for incorporating related environments using metadata collected during information gathering,

[0828] A means for editing and generating original video by combining the aforementioned analysis results and generated information,

[0829] A means of providing the aforementioned video to the user,

[0830] A means of providing the generated video in augmented reality format,

[0831] A system that includes this.

[0832] (Claim 2)

[0833] The system according to claim 1, wherein the aforementioned audio commentary is generated using technology that mimics the user's voice and is viewable in an augmented reality environment.

[0834] (Claim 3)

[0835] The system according to claim 1, wherein the aforementioned related environment is constructed using historical information based on the date and time of shooting and location information, and is incorporated into an augmented reality experience.

[0836] "Example 2 of combining an emotion engine"

[0837] (Claim 1)

[0838] A device for acquiring photographic and video data,

[0839] A device for analyzing the emotions of a subject based on the aforementioned data,

[0840] A device that generates music and narration based on the emotions of the subject and the emotional state of the user,

[0841] A device that incorporates related information by utilizing metadata acquired during data acquisition,

[0842] A device for editing and generating personalized videos by integrating the aforementioned analysis results and generated content,

[0843] A device that adjusts the tempo and atmosphere of the generated video based on the user's real-time emotional state,

[0844] A device that provides the aforementioned video to a user and makes it accessible to the user,

[0845] A system that includes this.

[0846] (Claim 2)

[0847] The system according to claim 1, wherein the narration is generated using a technique that imitates the user's voice.

[0848] (Claim 3)

[0849] The system according to claim 1, wherein the aforementioned related information is constructed using past data based on the date and time of shooting and location information.

[0850] "Application example 2 when combining with an emotional engine"

[0851] (Claim 1)

[0852] Means of acquiring digital information,

[0853] A means for analyzing the subject's emotions based on the aforementioned information,

[0854] A means for generating music and explanatory text that match the emotions of the subject,

[0855] A means of incorporating relevant background information by utilizing auxiliary information during data collection,

[0856] A means for editing and generating original video by combining the aforementioned analysis results and generated content,

[0857] A means of providing the aforementioned video to the user,

[0858] A means of collecting the user's emotional state in real time and adjusting the tempo and atmosphere within the video,

[0859] A system that includes this.

[0860] (Claim 2)

[0861] The system according to claim 1, wherein the explanatory text is generated using a technology that imitates the user's voice.

[0862] (Claim 3)

[0863] The system according to claim 1, wherein the aforementioned related background is constructed using temporal data based on the date and time of shooting and location information. [Explanation of Symbols]

[0864] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for acquiring photographic and video data, A means for analyzing the emotions of the subject based on the aforementioned data, Means for generating music and narration that match the emotions of the subject, A means of incorporating relevant context by utilizing metadata collected during data acquisition, A means for editing and generating an original video by combining the aforementioned analysis results and generated content, A means of providing the aforementioned video to the user, A system that includes this.

2. The system according to claim 1, wherein the narration is generated using technology that imitates the user's voice.

3. The system according to claim 1, wherein the aforementioned related context is constructed using historical data based on the date and time of shooting and location information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A