System

The system addresses the challenge of creating personalized movies by uploading, analyzing, and generating movies with user feedback, ensuring privacy and cultural adaptation, thus simplifying the creation and sharing process.

JP2026028729APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131345
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Existing systems fail to efficiently combine photos, videos, and audio recordings into a moving story, and lack comprehensive support for multilingualism, cultural adaptation, and privacy protection, making it cumbersome for users to create personalized movies.

Method used

A system that uploads, analyzes, and classifies user data, generates a storyboard based on preferences, adds music and effects, supports multiple languages and cultures, ensures privacy, and allows user feedback for corrections, ultimately exporting a personalized movie.

Benefits of technology

Enables users to easily create and share personalized movies with privacy protection, multilingual support, and cultural adaptation, enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028729000001_ABST
    Figure 2026028729000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system that includes means for uploading photo, video, and audio recording files, means for analyzing the uploaded data and extracting content categories and important scenes and messages, means for automatically generating impressive stories based on user desires and themes, means for adding music and effects to impressive stories, means for multilingual and cultural adaptation, means for presenting the generated movie to the user and making modifications based on feedback, means for ensuring privacy and data security, and means for exporting and sharing the final movie file in a specified format.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Today, there is an increasing demand for recording and sharing special events and life milestones. However, it is not easy to combine numerous photos, videos, and audio recordings into a moving story. Furthermore, analyzing this data and generating personalized movies based on user preferences requires advanced technology. Furthermore, requirements such as multilingual support, cultural adaptation, and privacy protection must also be met. Since no system exists that simultaneously meets these multiple requirements, this process can be cumbersome for users. Therefore, there is a need for a system that can automatically generate moving, personalized movies based on various data provided by users. [Means for solving the problem]

[0005] The present invention solves the above problems by providing a system that includes means for uploading photo, video, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, means for automatically generating an inspiring story based on the user's wishes and themes, means for adding music and effects to the inspiring story, means for multilingual support and cultural adaptation, means for presenting the generated movie to the user and making corrections based on feedback, means for ensuring privacy and data security, and means for exporting and sharing the final movie file in a specified format.

[0006] The system classifies data uploaded by users and tags key moments using facial recognition and scene detection technologies. It also uses speech recognition technology to convert conversation recordings into text and extract important messages and keywords. Based on this data, an AI model generates an inspiring storyboard and adds music and effects to create a visually appealing and moving movie. Furthermore, the system supports multiple languages ​​and cultural adaptations, providing translation and cultural customization to meet the needs of users and viewers. The final movie is then exported after users preview and provide feedback. Any necessary revisions are incorporated into the final version. Privacy and data security are given top priority when handling data, and encryption technology is used to protect it. This allows users to easily create and share personalized movies.

[0007] A "photograph" is a recorded still image, a means of visually representing a particular moment or scene.

[0008] "Video" is a recording of moving images, a means of providing audio and visual information simultaneously.

[0009] A "voice recording file" is a digital copy of audio data stored in a computer system and is a means of recording audio information, including conversations and messages.

[0010] "Upload" refers to the act of sending data from a local device to a remote server or cloud system.

[0011] "Data analysis" refers to the process of processing digital data, such as provided photographs, videos, and audio recordings, to understand and classify their content.

[0012] "Content classification" refers to organizing analyzed data into specific categories or themes.

[0013] A "significant scene" is a particularly meaningful moment or event in a photo or video that plays an important role as part of a larger story.

[0014] A "message" is meaningful text information extracted from a voice recording file, including personal words and expressions.

[0015] "Auto-generated storytelling" refers to the process by which an AI model creates a compelling, personalized narrative based on user-provided data.

[0016] "Music" refers to the acoustic elements generated to enhance emotion or atmosphere, and refers to songs added to movies.

[0017] "Effects" refers to technical techniques used to enhance the appeal of content by adding visual or auditory effects.

[0018] "Multilingual support" means providing subtitles and narration in multiple languages ​​to accommodate users with different language environments.

[0019] "Cultural adaptation" refers to adapting content to different cultures and customs, and providing culturally appropriate formats and expressions.

[0020] "Preview" refers to an operation in which a user checks the first version of a generated movie in advance and checks its contents.

[0021] "Feedback" refers to the process by which users communicate requests for corrections and improvements to the system based on the preview.

[0022] "Privacy protection" refers to measures to protect users' personal information and digital data from unauthorized access and information leaks.

[0023] "Data security" refers to the technologies and processes used to ensure the confidentiality, integrity, and availability of digital data.

[0024] "Export" refers to the process of outputting the generated movie file in a specific format so that it can be saved or used on another platform. [Brief explanation of the drawings]

[0025] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0026] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0027] First, the terms used in the following description will be explained.

[0028] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0029] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0030] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0031] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0033] [First embodiment]

[0034] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0035] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0036] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0037] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0038] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0040] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0041] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0042] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0043] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0044] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0045] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0046] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, it is necessary to design and implement programs based on the following means.

[0047] 1. Data collection

[0048] Users upload photos, videos, and audio files to the system through dedicated applications or websites. The digital data provided is then sent to the server.

[0049] 2. Data Analysis and Classification

[0050] Server: Receives uploaded data and categorizes photos, videos, and audio recording files into their respective categories. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are converted to text using speech recognition technology to extract important messages and keywords.

[0051] 3. Build a story

[0052] Server: Based on the user's specified theme and wishes, the server automatically generates the optimal story based on key moments and messages. The AI ​​model builds the storyboard and determines the order and narrative of each scene.

[0053] 4. Music and Effects Selection

[0054] Server: Based on the story, select music from the library to enhance emotions and add effects appropriate for each scene. The music and effects are adjusted to fit the flow of the scene.

[0055] 5. Multilingual and culturally accommodating

[0056] Server: Performs multilingual text translation to generate subtitles and narration in the user's language, and customizes content based on the user's culture and customs.

[0057] 6. Privacy and Data Security

[0058] Server: Uploaded digital data is protected with the latest encryption technology, and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[0059] 7. Generate and preview the movie

[0060] Server: Integrates all elements and generates the final movie file, which is temporarily stored and made available for user preview.

[0061] User: Can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[0062] 8. Final output and sharing

[0063] Server: Generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g. high-resolution video file or streaming link) can be selected based on the user's preference.

[0064] Device: Receive the final movie file and save it locally or share it on an online platform.

[0065] Specific examples

[0066] For example, consider the case of creating a wedding movie.

[0067] 1. Data collection

[0068] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[0069] 2. Data Analysis and Classification

[0070] Server: Analyzes the bride and groom's appearance in photos and tags the moments of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes recorded conversations to extract moving messages.

[0071] 3. Build a story

[0072] Server: Arrange individual photos and video scenes on a storyboard to build a moving story from the bride and groom meeting to their wedding.

[0073] 4. Music and Effects Selection

[0074] Server: Insert moving music and apply effects (fade in / out, slow motion, etc.) to suit the scene.

[0075] 5. Multilingual and culturally accommodating

[0076] Server: For international weddings, generate subtitles and narration translated into each country's language.

[0077] 6. Privacy and Data Security

[0078] Server: All data is encrypted and secured.

[0079] 7. Generate and preview the movie

[0080] User: Preview the generated movie and request changes to the scene order or music.

[0081] 8. Final output and sharing

[0082] Server: Generates the final version of the movie and serves it to the user.

[0083] Device: Download movies, save them, and share them online.

[0084] In this way, the system according to the present invention automatically generates an inspiring movie based on the digital data recording a special moment of the user, and provides it as a personalized keepsake.

[0085] The processing flow will be explained below.

[0086] Step 1:

[0087] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The data is temporarily stored on the device.

[0088] Step 2:

[0089] Terminal: The data uploaded by the user is sent to the server via the Internet, including metadata (such as date, time, and location information).

[0090] Step 3:

[0091] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[0092] Step 4:

[0093] Server: Analyzes photo data using image recognition technology, automatically tags faces and distinctive scenes (e.g., smiles, touching moments).

[0094] Step 5:

[0095] Server: The video analysis module analyzes the video data, detects scene changes, and extracts important moments, such as the engagement kiss and ring exchange.

[0096] Step 6:

[0097] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[0098] Step 7:

[0099] Server: Generates a storyboard using key scenes and messages based on the user's specified theme and wishes. The AI ​​model determines the order of scenes and narrative progression to build a compelling story.

[0100] Step 8:

[0101] Server: Select music from the library to enhance the emotions of the story and add effects to each scene, including fade-in / out and slow motion.

[0102] Step 9:

[0103] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content for cultural adaptation.

[0104] Step 10:

[0105] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0106] Step 11:

[0107] Server: Integrates all elements and generates the initial movie version. The generated movie is temporarily saved and a preview link is provided to the user.

[0108] Step 12:

[0109] Users: Use the preview link to see the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[0110] Step 13:

[0111] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[0112] Step 14:

[0113] Server: Generates the final version of your movie and exports it in the format of your choice, including high-resolution video files and streaming links.

[0114] Step 15:

[0115] Terminal: Receive the final movie file and save it locally. If you want to share it on an online platform, you can easily do so using a dedicated application.

[0116] These are the basic processing steps in the embodiment of the present invention, which allows users to effectively record special moments and easily share them as moving movies.

[0117] Example 1

[0118] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0119] There is a need for a system that allows users to easily create inspiring and personalized movies. It is also necessary to meet the needs of each user for privacy protection, data security, multilingual support, and cultural adaptation. Conventional systems cannot comprehensively provide these features, and they are time-consuming and labor-intensive.

[0120] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0121] In this invention, the server includes means for users to upload photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, means for automatically generating an inspiring story based on the user's wishes and themes, means for adding music and effects necessary for the inspiring story, means for providing multilingual support and cultural adaptation to the generated story, means for presenting the generated movie to the user and making corrections based on feedback, means for encrypting digital data to ensure privacy protection and data security, and means for exporting and sharing the final movie file in a specified format, thereby enabling users to easily create inspiring and personalized movies that also provide privacy protection, multilingual support, and cultural adaptation.

[0122] "Uploading" refers to the transfer of digital data from a user's terminal to a remote system such as a server.

[0123] "Data analysis" is the process of analyzing uploaded digital data, categorizing it into different formats such as photos, videos, and audio recording files, and understanding its content.

[0124] "Content classification" refers to the act of organizing analyzed digital data based on its type and characteristics, separating it into categories such as photos, videos, and audio recording files.

[0125] "Significant scene and message extraction" is the process of identifying and tagging particularly meaningful scenes and messages within photographs, videos, and audio recordings.

[0126] "Means for automated story generation" refers to algorithms and AI models that create compelling narrative structures from digital data based on user preferences and themes.

[0127] "Adding music and effects" is the process of applying emotionally enhancing music and visual effects to the generated story.

[0128] "Multilingual support" refers to the ability to translate generated content into multiple user-specified languages ​​and include subtitles and narration.

[0129] "Cultural adaptation" is the process of appropriately customizing content based on the user's specified culture and customs.

[0130] "Privacy protection" refers to various measures taken to protect users' personal information and digital data from others.

[0131] "Data security" means the technological measures and processes used to ensure the confidentiality, integrity and protection from unauthorized access of uploaded digital data.

[0132] "Export" refers to the act of converting the final generated movie file into a specific format (e.g., MP4, AVI, etc.) and outputting it.

[0133] "Sharing" refers to providing the generated movie file to other users or platforms so that they can access it.

[0134] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, the following hardware and software are used:

[0135] 1. Hardware and Software

[0136] Hardware: Servers, user devices (PCs, smartphones, etc.)

[0137] Software: Uploading applications and websites, image recognition technology (e.g., OpenCV, TensorFlow), speech recognition technology (e.g., Google Speech-to-Text API, IBM Watson), multilingual text translation systems (e.g., Google Translate API, DeepL), AI models (e.g., GPT-4)

[0138] 2. Detailed processing instructions

[0139] 1. Data collection

[0140] Users upload digital data (e.g., photos, videos, and audio files from the wedding day) through a dedicated application or website, and this data is sent to the server.

[0141] 2. Data Analysis and Classification

[0142] The server receives the uploaded data and categorizes it into photos, videos, and audio recordings. Image recognition technology is used to identify scenes in the photos and videos, confirm the bride and groom's appearance, and tag moments such as the vow kiss and ring exchange. Audio files are converted to text using speech recognition technology, which extracts important messages and keywords.

[0143] 3. Build a story

[0144] The server automatically generates the optimal story based on the user's specified theme and wishes, key moments, and messages. It uses a generative AI model (e.g., GPT-4) to build the storyboard and determine the order and narrative of each scene.

[0145] 4. Music and Effects Selection

[0146] The server selects inspiring music from a library and adds appropriate effects for each scene, with the music and effects tailored to the flow of the scene.

[0147] 5. Multilingual and culturally accommodating

[0148] The server uses a multilingual translation system to generate subtitles and narration in the language specified by the user, and also customizes the content based on the specified culture and customs.

[0149] 6. Privacy and Data Security

[0150] The server protects uploaded digital data using the latest encryption technology (e.g., AES-256), and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[0151] 7. Generate and preview the movie

[0152] The server aggregates all the elements and generates the final movie file using a video editing library (e.g., FFmpeg), which is temporarily stored in cloud storage and made available for users to preview.

[0153] 8. Final output and sharing

[0154] The server generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g., high-resolution video file or streaming link) can be selected based on the user's preference. The user's device receives the final movie file and can save it locally or share it on an online platform.

[0155] 3. Examples of concrete examples and prompts

[0156] A specific example of creating a wedding movie will be given below.

[0157] Data collection

[0158] Users upload photos and videos from their wedding day, as well as recorded messages from friends and family.

[0159] Data analysis and classification

[0160] The server analyzes the bride and groom's appearance in photos and tags the moment of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes audio messages to extract moving messages.

[0161] Building a story

[0162] The server uses these photos and video footage to build a moving story from the moment the bride and groom met to their wedding day.

[0163] Music and effects selection

[0164] Choose inspiring music and add effects that suit the scene (e.g. fade in / out, slow motion, etc.).

[0165] Multilingual and culturally accommodating

[0166] For international weddings, generate subtitles and narration translated into each country's language.

[0167] Privacy and Data Security

[0168] Encrypt your data to ensure the security of all your data.

[0169] Generate and preview the movie

[0170] The user can preview the generated movie and request changes to the scene order or music.

[0171] Final output and sharing

[0172] The server generates the final version of the movie and serves it to the user, who then downloads it, saves it, and shares it online.

[0173] Example prompt sentence:

[0174] "I would like to create a wedding video using the bride and groom's kiss and ring exchange scenes from the photos and videos I uploaded, and create a story with inspiring music. I would like the narration to be in both English and Spanish."

[0175] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0176] Step 1:

[0177] Data collection

[0178] Input: Photos, videos, and audio files from the user's wedding day

[0179] The user selects the digital data using a dedicated application or website and clicks an upload button.

[0180] The terminal transmits the selected data to a server via the Internet.

[0181] Output: Digital data sent to the server

[0182] Step 2:

[0183] Data analysis and classification

[0184] Input: Digital data sent to the server

[0185] The server runs a data analysis program that categorizes the photos, videos, and audio recordings into their respective categories.

[0186] The server uses image recognition technology (e.g., OpenCV, TensorFlow) to analyze the photos and video scenes and tag important moments such as the bride and groom's faces, the ring exchange, and the vow kiss.

[0187] The server uses voice recognition technology (e.g., Google Speech-to-Text API, IBM Watson) to convert the audio recording file into text and extract inspirational messages and keywords.

[0188] Output: Classified digital data, tagged key scenes, and audio data converted to text

[0189] Step 3:

[0190] Building a story

[0191] Input: Classified digital data, tagged important scenes, transcribed audio data, user preferences and themes

[0192] The server generates a storyboard using an AI model (e.g., GPT-4) based on the theme and wishes (prompt sentences) specified by the user.

[0193] The server places the tagged scenes and extracted messages on a storyboard and determines the order of each scene.

[0194] Output: Generated storyboard

[0195] Step 4:

[0196] Music and effects selection

[0197] Input: Generated storyboard

[0198] The server selects suitable songs from a music library (e.g., Epidemic Sound, Artlist) to enhance the emotion.

[0199] The server applies visual effects such as fade-in, fade-out, and slow motion to match the storyboard scenes.

[0200] Output: Story with added music and effects

[0201] Step 5:

[0202] Multilingual and culturally accommodating

[0203] Input: Stories with added music and effects, user-specified language and cultural preferences

[0204] The server uses a multilingual translation system (e.g., Google Translate API, DeepL) to translate subtitles and narration into the language specified by the user.

[0205] The server customizes the story content based on the specified culture and customs.

[0206] Output: Multilingual and culturally adapted stories

[0207] Step 6:

[0208] Privacy and Data Security

[0209] Input: Uploaded digital data

[0210] The server encrypts all digital data using the latest encryption technology (e.g., AES-256).

[0211] The server uses AI models to anonymize and mask unnecessary parts of the data to protect user privacy.

[0212] Output: Encrypted and privacy-protected digital data

[0213] Step 7:

[0214] Generate and preview the movie

[0215] Input: Multilingual and culturally adapted stories, encrypted and privacy-protected digital data

[0216] The server uses a video editing library (e.g. FFmpeg) to stitch all the elements together and generate the final movie file.

[0217] The server temporarily stores the generated movie file in cloud storage and provides a preview page URL for the user to preview it.

[0218] Users can visit a preview page to view the generated movie and provide feedback through the interface on any necessary corrections.

[0219] Output: Movie file with user review and feedback

[0220] Step 8:

[0221] Final output and sharing

[0222] Input: Movie file reviewed and feedback by user, user preferred format

[0223] The server makes corrections based on the user's feedback and generates the final version of the movie.

[0224] The server outputs the movie file in the user's desired export format (e.g., high-resolution video file, streaming link, etc.).

[0225] The user's device downloads the final movie file and stores it locally, allowing them to share it online via social media or email if desired.

[0226] Output: Final version movie file

[0227] (Application example 1)

[0228] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0229] This invention relates to a system that automatically generates moving, personalized movies simply by providing digital data (photos, videos, and audio recording files) from users. However, conventional systems lack specific application examples for further improving the user experience. In particular, they are unable to record shopping experiences in virtual stores and automatically generate movies from them, making it difficult for users to experience moving movies in different contexts. To solve this problem, it is necessary to expand the system to include movie generation based on shopping experiences in virtual stores.

[0230] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0231] In this invention, the server includes a means for uploading photos, videos, and audio recording files, a means for analyzing the uploaded data, categorizing the content, and extracting important scenes and messages, and a means for automatically generating an inspiring story based on the user's preferences and themes. This enables a shopping experience movie to be automatically generated using data taken by the user in a virtual store. The server also includes a means for performing facial recognition and tagging important moments, a means for tagging products and scenes in the virtual store, a means for converting conversation recording files into text using voice recognition technology and extracting important messages and keywords, and a means for generating inspiring prompts based on the user's shopping experience in the virtual store, further enriching the user's shopping experience and recording it as an inspiring movie.

[0232] "Photography" is a visual medium that uses light to record the shape and color of objects.

[0233] "Video" is a medium containing visual and audio information that records a time-sequence of images.

[0234] An "audio recording file" is a file that digitally records a person's voice or sound.

[0235] "Uploading" refers to the act of sending data owned by a user to a server via the Internet.

[0236] "Analysis" is the act of breaking down the provided data to understand its content and structure.

[0237] "Classification" is the act of dividing data into groups according to specific criteria.

[0238] An "important scene" is a moment that has particular meaning or impact in the overall flow or story.

[0239] A "message" is a portion of a voice recording that contains meaning or information extracted from the voice recording file.

[0240] "Automatic story generation" refers to the act of creating a coherent narrative based on analyzed data, based on the user's wishes and themes.

[0241] "Music" is sound with a melody or rhythm used to add emotional elements to a movie.

[0242] "Effects" are techniques and methods for giving a movie special visual or auditory effects.

[0243] "Multilingual support" refers to the ability to provide content in multiple languages.

[0244] "Cultural adaptation" is the act of adjusting content to suit a particular culture and customs.

[0245] "Privacy protection" refers to measures to protect personal information from unauthorized use.

[0246] "Data security" refers to the state in which data is protected from unauthorized access and tampering.

[0247] "Exporting" is the act of outputting the final movie file in a particular format.

[0248] "Sharing" is the act of making the generated movie available for viewing by others.

[0249] A "virtual store" is a virtual sales environment that offers products and services over the Internet.

[0250] "Shopping experience" is a general term for a series of actions and emotions that a user goes through when selecting and purchasing a product.

[0251] A "prompt sentence" is a sentence that is input to a generative AI model based on specified conditions and content.

[0252] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this invention, a system based on the following means is required.

[0253] 1. Data collection

[0254] The server provides a means for users to upload photos, videos, and audio recording files through dedicated applications or websites, making it easy for users to provide data.

[0255] 2. Data Analysis and Classification

[0256] The server receives the uploaded data and categorizes the photos, videos, and audio recordings into their respective categories. Image recognition technology is used to identify scenes in the photos and videos and tag specific moments. Audio recordings are converted into text using speech recognition technology, and important messages and keywords are extracted. Software such as OpenCV and TensorFlow are used for this process.

[0257] 3. Build a story

[0258] The server automatically generates the best story based on the user's specified theme and desires, key moments, and messages. It uses a generative AI model to build a storyboard and determine the order and narrative of each scene, creating a coherent and moving story.

[0259] 4. Music and Effects Selection

[0260] The server selects music from a library to enhance emotions based on the story, and adds effects appropriate for each scene. The music and effects are adjusted to match the flow of the scenes. This process is performed using software such as PIL (Python Imaging Library) and moviepy.

[0261] 5. Multilingual and culturally accommodating

[0262] The server performs multilingual text translation to generate subtitles and narration in the user's language, and customizes the content based on the user's culture and customs. By using translation software such as Google Translate API, it can accommodate users from different cultures.

[0263] 6. Privacy and Data Security

[0264] The server protects uploaded digital data using the latest encryption technology. AI models also anonymize data and mask unnecessary parts to maximize user privacy. Security software such as AES (Advanced Encryption Standard) is used.

[0265] 7. Generate and preview the movie

[0266] The server combines all the elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. Users can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[0267] 8. Final output and sharing

[0268] The server generates the final version of the movie incorporating the user's feedback and corrections. The user can choose the export format of the movie file (e.g., high-resolution video file or streaming link) according to their preference. The user receives the final movie file and can save it locally or share it on an online platform.

[0269] As a concrete example, consider a scenario where a user uploads photos and videos of clothes they tried on in a virtual store, and a shopping experience video is generated based on those photos and videos. Here is an example of a prompt to be input to the generative AI model:

[0270] "Create an inspiring shopping experience video using photos and videos of users trying on different outfits in a virtual store. At the end of the video, highlight the user's favorite outfit and play inspiring music in the background."

[0271] This prompt allows the generative AI model to generate a user experience movie that can then be previewed and shared.

[0272] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0273] Step 1:

[0274] Users upload photos, videos, and audio files through dedicated applications or websites. The input is the user's digital data, and the output is data stored on the server. Specifically, the user selects the captured data and presses the upload button, which sends the file to the server.

[0275] Step 2:

[0276] The server analyzes the uploaded data and classifies photos, videos, and audio recording files into their respective categories. The input is the uploaded digital data, and the output is the classified data. Specifically, image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are also converted into text using speech recognition technology, and important messages and keywords are extracted. This processing is done using OpenCV and TensorFlow.

[0277] Step 3:

[0278] The server references the user's specified theme and wishes, and automatically generates an optimal story based on key moments and messages. The input is classified data and the user's theme settings, and the output is a storyboard. Specifically, a generative AI model analyzes the data and constructs the story flow. This process generates a coherent and moving story.

[0279] Step 4:

[0280] Based on the story, the server selects music from a library to enhance emotions and adds appropriate effects to each scene. The input is a storyboard, and the output is the final video scene. Specifically, PIL (Python Imaging Library) and moviepy are used to insert music and effects into the video. This process creates visually and aurally appealing content.

[0281] Step 5:

[0282] The server performs multilingual text translation to generate subtitles and narration in the user's specified language. It also customizes the content based on the specified culture and customs. The input is the story text information and the user's specified language, and the output is the translated subtitles and narration. Specifically, it uses the Google Translate API or similar to translate the text data and generate the required subtitles and narration.

[0283] Step 6:

[0284] The server protects uploaded digital data using the latest encryption technology. Furthermore, the AI ​​model anonymizes the data and masks unnecessary parts to ensure maximum user privacy. The input is the user's digital data, and the output is protected data. Specifically, the data is encrypted using security technologies such as AES (Advanced Encryption Standard).

[0285] Step 7:

[0286] The server aggregates all elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. The input is the aggregated content, and the output is the generated movie file. Specifically, software such as moviepy is used to combine all scenes and effects and create the final movie file.

[0287] Step 8:

[0288] The user previews the generated movie and provides feedback through the interface on any necessary modifications (e.g., changing the order of scenes or replacing music). The input is the generated movie file and the user's feedback, and the output is the modified movie file. Specifically, the server performs the modifications based on the feedback provided by the user.

[0289] Step 9:

[0290] The server generates the final version of the movie incorporating the modifications based on the user's feedback. The export format (e.g., high-resolution video file or streaming link) can be selected according to the user's preference. The input is the modified movie file, and the output is the final movie file. Specifically, we provide the function to export the movie file in the format desired by the user.

[0291] Step 10:

[0292] Users receive the final movie file and can save it locally or share it on an online platform. The input is the final movie file, and the output is the shared movie. Specifically, the generated movie file can be uploaded to social media or cloud storage to be shared with other users.

[0293] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0294] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by the user. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more emotionally rich stories. To implement the system, it is necessary to design and implement programs based on the following means.

[0295] 1. Data collection

[0296] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. These data are temporarily stored on the device.

[0297] 2. Data transmission

[0298] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[0299] 3. Data Analysis and Classification

[0300] Server: Receives uploaded data and classifies photos, videos, and audio recordings. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recordings are also converted to text using speech recognition technology to extract important messages and keywords.

[0301] 4. Emotional Recognition

[0302] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[0303] 5. Build a story

[0304] Server: Based on the user's specified themes and wishes, and taking into account the emotional information recognized by the emotion engine, the server automatically generates a moving story. The AI ​​model determines the appropriate scene order and narrative to build a compelling story.

[0305] 6. Music and Effects Selection

[0306] Server: Based on the story, select music from the library to enhance the emotion and add effects to suit the scene, such as fade in / out and slow motion.

[0307] 7. Multilingual and culturally adaptable

[0308] Server: Generates subtitles and narration based on the user-specified language. Uses text translation tools for multilingual support and adjusts content for cultural adaptation.

[0309] 8. Privacy and Data Security

[0310] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0311] 9. Generate and preview the movie

[0312] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and users can view it using the preview link.

[0313] Users: can use the preview link to see the generated movie and provide feedback, such as requests for changes to the scene order or music.

[0314] 10. Receiving and Responding to Feedback

[0315] Server: Receives user feedback and modifies the movie based on the instructions. The modified movie is then regenerated and saved.

[0316] 11. Final output and sharing

[0317] Server: Generates the final version of the movie and provides it to the user in the export format of their choice, such as a high-resolution video file or a streaming link.

[0318] Device: Receive the final movie file and save it locally or share it on an online platform.

[0319] Specific examples

[0320] For example, consider the case of creating a wedding movie.

[0321] 1. Data collection

[0322] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[0323] 2. Data transmission

[0324] Terminal: Sends these uploaded data to the server.

[0325] 3. Data Analysis and Classification

[0326] Server: Recognizes and tags important moments in photos and videos (e.g., the promise kiss, the ring exchange), and extracts inspirational messages from audio recordings.

[0327] 4. Emotional Recognition

[0328] Server: Analyzes facial expressions and tone of voice from video and audio to recognize the emotions of the bride and groom and guests. For example, it identifies touching moments by detecting smiles and tears.

[0329] 5. Build a story

[0330] Server: Based on the extracted scenes and messages and the emotional information recognized by the emotion engine, a moving story of the entire wedding is constructed.

[0331] 6. Music and Effects Selection

[0332] Server: Add inspiring music and apply effects appropriate to the scene, including fade in / out, slow motion, etc.

[0333] 7. Multilingual and culturally adaptable

[0334] Server: For international weddings, generate subtitles and narration translated into each country's language.

[0335] 8. Privacy and Data Security

[0336] Server: All data is encrypted and secured.

[0337] 9. Generate and preview the movie

[0338] User: Preview the generated movie and request corrections if necessary.

[0339] 10. Receiving and Responding to Feedback

[0340] Server: Based on user feedback, the movie is regenerated and modifications are made.

[0341] 11. Final output and sharing

[0342] Server: Generates the final version of the movie and serves it to the user.

[0343] Device: Download movies and finally save and share them.

[0344] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[0345] The processing flow will be explained below.

[0346] Step 1:

[0347] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. Uploaded data is temporarily stored on the device.

[0348] Step 2:

[0349] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[0350] Step 3:

[0351] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[0352] Step 4:

[0353] Server: Analyzes photo data using image recognition technology, automatically identifying faces and tagging distinctive scenes (e.g., smiles or touching moments).

[0354] Step 5:

[0355] Server: Analyzes video data using the video analysis module, detects scene changes, and extracts important moments (e.g., the engagement kiss or ring exchange).

[0356] Step 6:

[0357] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[0358] Step 7:

[0359] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[0360] Step 8:

[0361] Server: Automatically generates an emotional storyboard based on the user's specified themes and wishes, and also takes into account recognized emotional information. The AI ​​model determines the order of scenes and the progression of the story, building a coherent story.

[0362] Step 9:

[0363] Server: Based on the story, select music from the library to enhance the emotions and add effects to each scene, such as fade in / out and slow motion.

[0364] Step 10:

[0365] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content to take cultural adaptation into account.

[0366] Step 11:

[0367] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0368] Step 12:

[0369] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and a preview link is provided to users.

[0370] Step 13:

[0371] Users: Use the preview link to view the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[0372] Step 14:

[0373] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[0374] Step 15:

[0375] Server: Generates the final movie file and exports it in the format of your choice (e.g. high-resolution video file or streaming link).

[0376] Step 16:

[0377] Terminal: Receive the final movie file and save it locally. Furthermore, if you want to share it on an online platform, you can easily do so using a dedicated application.

[0378] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[0379] Example 2

[0380] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0381] In recent years, the widespread use of smartphones and digital cameras has increased the opportunities for users to save various events and everyday moments as digital data (photos, videos, and audio recording files). However, editing this data to create moving and personalized movies requires advanced editing skills and time, making it difficult for average users. Conventional methods have been problematic in that the editing process is cumbersome and it is difficult to automatically generate emotionally appealing stories, leaving users unable to obtain satisfactory results.

[0382] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to upload photos, videos, and audio recording files; a means for transmitting the uploaded data to the server via the Internet; a means for the server to analyze the received data, classify the content, and extract important scenes and messages; a means for analyzing the user's emotions using emotion recognition technology; a means for automatically generating a story based on the user's preferences and themes, taking into account the emotional information; a means for adding music and effects to the story; a means for multilingual support and cultural adaptation; a means for presenting the generated movie to the user and making corrections based on the user's feedback; a means for ensuring privacy protection and data security; and a means for exporting and sharing the final movie file in a specified format. This allows even ordinary users to easily automatically generate moving and personalized movies and obtain high-quality results.

[0383] "User" refers to any individual or organization that uses this system.

[0384] "Photo" refers to still image data that a user takes using a digital device and uploads to the system.

[0385] "Video" refers to moving image data that a user shoots using a digital device and uploads to the system.

[0386] "Audio recording file" refers to a file that contains audio data recorded by a user or another person.

[0387] "Upload" refers to a user sending a photo, video, or audio recording file to the system through a dedicated application or website.

[0388] "Internet" refers to a global network for data communication.

[0389] "Server" refers to a computer system that receives, analyzes, sorts, and processes data sent by users.

[0390] "Analysis" refers to the process by which the server understands and interprets the content of photographs, videos, and audio recordings based on data.

[0391] "Classification" refers to the process by which the server divides the data it receives into specific categories.

[0392] "Important scenes and messages" refer to notable moments or words that the user finds interesting or that the system automatically selects.

[0393] "Emotion recognition technology" refers to the technology that a system uses to analyze and identify a user's emotions from images and audio.

[0394] "Automatic story generation" refers to the process in which a system constructs a series of events into a narrative based on the user's wishes, themes, and emotional information.

[0395] "Music and Effects" refers to background music and visual effects added to enhance the visual and auditory impression of the movie.

[0396] "Multilingual" refers to the ability to display content or provide narration in different languages.

[0397] "Cultural adaptation" refers to adjusting content to suit the user's culture.

[0398] "Feedback" refers to the act of a user conveying to the system their opinions and requests for corrections and improvements to a movie.

[0399] "Privacy protection" refers to protecting users' personal information and data from being leaked to third parties.

[0400] "Data security" refers to measures taken to keep data protected from unauthorized access and tampering.

[0401] "Export" refers to the act of saving or outputting the final movie file in a particular format.

[0402] "Sharing" refers to sharing the generated movie file or link with other users and platforms.

[0403] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. Specific procedures for implementing the present invention and the hardware and software used are described below.

[0404] This system is primarily composed of users, devices, and a server. Users upload data using a dedicated application or website, and their devices send the data to the server via the Internet. The server analyzes the data, generates a story, and provides the final movie file to the user. The specific operation is explained below.

[0405] Data collection

[0406] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website, and these data are temporarily stored on the device.

[0407] Sending data

[0408] The device sends the data uploaded by the user to a server via the Internet, including metadata (e.g., date, time, location, etc.).

[0409] Data analysis and classification

[0410] The server analyzes the received data and classifies each photo, video, and audio recording file using image recognition technologies (e.g., Google Cloud Vision API) and speech recognition technologies (e.g., Google Cloud Speech-to-Text). Scenes in the photos and videos are identified and specific moments (e.g., the engagement kiss or the ring exchange) are tagged. Audio recordings are converted to text and key messages and keywords are extracted.

[0411] Emotion recognition

[0412] The server uses emotion recognition technology (e.g., Microsoft Azure Emotion API) to recognize the user's emotions from video and audio, analyzing smiles, tears, tone of voice, etc. to identify emotional moments.

[0413] Building a story

[0414] The server automatically generates a moving story based on the user's preferences and themes, taking into account emotional information. This story generation uses an AI model (e.g., OpenAI GPT) to determine the order of movie scenes and the content of the narration.

[0415] Music and effects selection

[0416] Based on the story, the server selects music from a library (e.g., Epidemic Sound) to enhance the emotions and adds effects (fade in / out, slow motion, etc.) that match the scene.

[0417] Multilingual and culturally accommodating

[0418] The server generates subtitles and narration based on the language specified by the user, and uses text translation tools (e.g., Google Translate API) to support multiple languages, adjusting the content for cultural adaptation.

[0419] Privacy and Data Security

[0420] The server uses the latest encryption technology (e.g., AES-256) for all uploaded data to ensure privacy and data security. Data is sent and received using SSL / TLS encryption, and data is encrypted when stored.

[0421] Generate and preview the movie

[0422] The server combines all the elements to generate a first version of the movie file, stores it temporarily, and makes it available to users via a preview link.

[0423] Users can use the preview link to view the resulting movie and submit requests for changes to the scene order or music, if necessary.

[0424] Receiving and responding to feedback

[0425] The server receives feedback from the user and modifies the movie based on the instructions, and the modified movie is regenerated and saved.

[0426] Final output and sharing

[0427] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link.

[0428] The device will receive the final movie file and store it locally, or you can share it on an online platform.

[0429] Examples of prompt statements

[0430] Below are some example prompts to input to the generative AI model:

[0431] Create a moving wedding video using photos and videos from your wedding, along with messages from friends and family. Highlight the bride and groom's smiling and tearful moments, and add music appropriate to their special day. For an international wedding, add subtitles in English and Japanese to reflect cultural elements. Specifically, use emotional music for the groom's entrance and select a song that gradually builds up during the bride's entrance. Additionally, highlight the vows and ring exchange in slow motion, and add moving guest messages.

[0432] Through the above steps, the present invention utilizes an emotion engine to automatically generate an even more moving movie based on digital data that records a user's special moments, and provides it as a personalized keepsake.

[0433] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0434] Step 1: Collect data

[0435] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The input is the user's own digital data. The uploaded data is temporarily stored on the device and becomes the output. For example, a user can drag and drop photos from their wedding day or audio files containing messages from friends into the application, and then enter the event theme and desired details into a form.

[0436] Step 2: Sending data

[0437] The device sends the data uploaded by the user to the server via the Internet. The input includes data temporarily stored on the device and metadata (date, time, location, etc.). The output is data sent to the server and received by the server. Specifically, the device securely sends this data to the server using the HTTPS protocol.

[0438] Step 3: Analyze and classify the data

[0439] The server analyzes and classifies the data it receives. The input is the data received by the server. The output is photos, videos, and audio recordings sorted into their respective categories, with each file tagged with important scenes and keywords. For example, the server uses Google Cloud Vision API to identify scenes in photos and videos and tag specific moments, such as the bride and groom's kiss and ring exchange. It also uses Google Cloud Speech-to-Text to convert audio recordings into text and extract important messages and keywords.

[0440] Step 4: Recognize emotions

[0441] The server uses emotion recognition technology to analyze user emotions from video and audio. The input is analyzed and classified video and audio data. The output is an emotional tag attached to each piece of data. Specifically, it uses the Microsoft Azure Emotion API to analyze smiles, tears, tone of voice, and other emotions to identify moments when the bride and groom and guests are emotional.

[0442] Step 5: Build your story

[0443] The server automatically generates a story based on the user's wishes and themes, taking into account emotional information. The inputs are data tagged with emotions and the wishes and themes provided by the user. The output is a movie plan that determines the order of scenes and narration for a moving story. Specifically, it uses OpenAI GPT to construct the story's narrative and determine the placement of scenes.

[0444] Step 6: Choose music and effects

[0445] The server selects music and effects based on the story. The input is a story plan. The output is a story plan with music and effects applied to scenes to enhance the emotions. Specific operations include selecting appropriate music from the Epidemic Sound library and applying effects such as fade-in / out and slow motion to scenes.

[0446] Step 7: Multilingualism and cultural adaptation

[0447] The server generates subtitles and narration based on the language specified by the user. The inputs are the language information specified by the user and a story plan. The output is a story plan that includes subtitles and narration in multiple languages. Specifically, the system uses the Google Translate API to translate the text into multiple languages ​​and edits the content to take cultural adaptation into account.

[0448] Step 8: Privacy and Data Security

[0449] The server uses encryption technology when sending, receiving, and storing data. The inputs include user data stored on the server and data being sent and received. The output is encrypted data. Specifically, data security is ensured using AES-256 data encryption and SSL / TLS protocols.

[0450] Step 9: Generate and preview your movie

[0451] The server integrates all elements to generate a first-run movie file and temporarily stores it. The input is the final story plan. The output is the first-run movie file. Specifically, video editing software is used to generate the movie based on the pre-planned layout.

[0452] The user uses the preview link to view the generated movie. The input is the generated movie file. The output is feedback based on the preview results.

[0453] Step 10: Receiving and responding to feedback

[0454] The server receives user feedback and modifies the movie based on the instructions. The inputs are the user's feedback and the original movie file. The output is a modified movie file that reflects the feedback. Specifically, the server reorders the requested scenes and adjusts the music to generate a new movie.

[0455] Step 11: Final output and sharing

[0456] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link. The input is the final, modified movie plan. The output is the final movie file. Specific operations include generating an HD MP4 file or a YouTube streaming link as the final output.

[0457] The device receives the final movie file and stores it locally. The input is the final movie file provided by the server. The output is a locally stored movie file, which the user can also upload to an online platform for sharing.

[0458] (Application example 2)

[0459] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0460] Conventional digital content generation systems that use photos, videos, and audio files have struggled to automatically generate emotionally rich stories. Furthermore, there are challenges, such as generating movies that effectively reflect the user's emotions, supporting multiple languages, and protecting privacy. In particular, there is a need for a system that can visually and emotionally recreate users' memories and share them in high resolution with simple smartphone application operations.

[0461] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, and means for automatically generating an emotional story based on the user's wishes and themes. This makes it possible to automatically generate an emotional digital story that effectively reflects the user's emotions using an emotion engine.

[0462] A "photograph" is a file that digitally records a still image.

[0463] A "video" is a file that contains a digital recording of moving images.

[0464] An "audio recording file" is a file that records audio in digital format.

[0465] "Uploading" means sending local data owned by the user to a server via the Internet.

[0466] "Data analysis" refers to the use of machine learning and algorithms to process uploaded digital content and extract specific information.

[0467] "Content classification" refers to grouping photos, videos, and audio recording files based on their format and characteristics.

[0468] "Key Scene and Message Extraction" refers to automatically identifying moments and meaningful messages within digital content that deserve special emphasis.

[0469] "User's wishes and themes" are requests and objectives for a story or movie specified by the user.

[0470] "Automatic generation of emotional stories" is the process of automatically constructing emotional stories based on extracted scenes and messages.

[0471] An "emotion engine" is an algorithm that recognizes user emotions from digital data and uses that information to create stories and movies.

[0472] "Adding Music and Effects" is the process of adding music and visual effects to a moving story to enhance its emotional expression.

[0473] "Multilingual support" means translating generated content into multiple languages ​​to accommodate users from different cultures and regions.

[0474] "Cultural adaptation" means adjusting content to suit users in a particular culture or region.

[0475] "Privacy protection" refers to measures and technologies to protect user data and information from third parties.

[0476] "Data security" means ensuring that digital data is protected and that unauthorized access is prevented.

[0477] "Final Movie File" is the finished digital movie file exported in the user's desired format.

[0478] "Export in specified format" means outputting the final movie in a file format selected by the user.

[0479] "Emotion analysis" refers to recognizing and identifying a user's emotional state from digital data.

[0480] "Previewing a movie" means checking a movie before it is completed.

[0481] "Accepting a change request" means receiving a request for correction or addition from a user and changing the content.

[0482] A "high resolution movie file" is a digital movie file with high quality video output.

[0483] "Share on social media" means posting and sharing the generated movie file on an internet communication platform.

[0484] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. By combining this system with an emotion engine, it creates a story that reflects the user's emotions.

[0485] Program Generation and Processing Description

[0486] Hardware and software used:

[0487] Server: High performance data processing and storage server.

[0488] Terminal (smartphone or tablet): A device through which a user uploads digital data.

[0489] Network Connection: An internet connection for uploading, downloading, and communicating data.

[0490] Software used:

[0491] TensorFlow: A machine learning framework for image and speech recognition.

[0492] OpenCV: A library for image and video analysis.

[0493] Django: A server-side web framework.

[0494] React Native: A development platform for smartphone applications.

[0495] Amazon Web Services (AWS) S3: Data storage service.

[0496] Amazon Polly: A speech synthesis service.

[0497] Examples of implementation:

[0498] 1. Uploading data:

[0499] Users use a smartphone app to upload photos, videos, and audio recordings related to events such as weddings and birthdays, and metadata (date, time, and location information) is automatically captured and added to the uploaded data.

[0500] 2. Data transmission:

[0501] The device sends the data uploaded by the user to a server via the Internet, where it is stored in Amazon Web Services (AWS) S3.

[0502] 3. Data analysis and classification:

[0503] The server receives uploaded data using the Django framework and leverages TensorFlow and OpenCV to analyze and classify photo, video, and audio recording files, automatically extracting and tagging important scenes and messages from the uploaded data.

[0504] 4. Emotion Recognition:

[0505] The server uses an emotion engine to analyze emotions in the data, for example by analyzing facial expressions and tone of voice in the video to identify moments of emotion or joy.

[0506] 5. Build a story:

[0507] The server automatically generates an emotional story based on the user's specified theme and wishes, and this process incorporates the emotional information recognized by the emotion engine.

[0508] 6. Add music and effects:

[0509] The server adds appropriate music and effects to the generated story, and also generates voice narration using Amazon Polly, including multilingual support and cultural adaptation.

[0510] 7. Generate and preview the movie:

[0511] The server generates the initial movie file, allows users to preview it through a smartphone app, and accepts and responds to user requests for changes.

[0512] 8. Final movie output:

[0513] Users can download the final version of their movie in high resolution and share it on social media, all while ensuring privacy and data security.

[0514] Examples of prompts:

[0515] "Analyze family trip data to create an inspiring movie that captures moments of happiness for users."

[0516] In this way, the system of the present invention utilizes digital data of special moments provided by the user to automatically generate moving and personalized movies, allowing memories to be recreated in a deeper and more emotional way.

[0517] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0518] Step 1:

[0519] Data collection and upload

[0520] Users use a smartphone application to upload photos, videos, and audio recordings. Specifically, they tap the app's "upload" button and select media files from their device. The input data are photos, videos, and audio recordings and their associated metadata. The output data is digital data sent to a server.

[0521] Step 2:

[0522] Sending data

[0523] The device sends the data uploaded by the user to a server via the Internet. Specifically, the application transfers the selected file to the server using the HTTPS protocol. The input data is the digital content specified by the user through the application. The output data is the file stored on the server.

[0524] Step 3:

[0525] Data analysis and classification

[0526] The server receives, analyzes, and classifies the uploaded data. It uses TensorFlow and OpenCV to analyze photo, video, and audio recording files and processes the data according to their respective formats. Specific operations include scene analysis using image recognition and text conversion using voice recognition. The input data is the uploaded digital data. The output data is the classified content and its metadata.

[0527] Step 4:

[0528] Emotion recognition

[0529] The server uses an emotion engine to analyze users' emotions from uploaded videos and audio. Specifically, it analyzes facial expressions and tone of voice to identify emotional elements and emotional moments. The input data are classified photos, videos, and audio recordings. The output data is media data with recognized emotion information.

[0530] Step 5:

[0531] Building a story

[0532] The server automatically generates an inspiring story based on the user's specified theme and wishes. It primarily uses an AI model to construct the story based on recognized emotional information and tagged scenes. Specifically, it determines the order of scenes and automatically forms a narrative. The input data is media data with emotional information and the user's specified theme. The output data is the constructed story.

[0533] Step 6:

[0534] Adding music and effects

[0535] The server adds appropriate music and visual effects to the generated story. It uses Amazon Polly to generate narration, and also supports multiple languages ​​and cultural adaptations. Specific operations include selecting music, applying effects, and inserting narration. The input data is the constructed story. The output data is a movie with music and effects added.

[0536] Step 7:

[0537] Generate and preview the movie

[0538] The server generates the final movie file and allows users to preview it through a smartphone app. Specific operations include rendering the movie file and generating a preview file. The input data is movie data with music and effects added. The output data is a movie file that users can preview.

[0539] Step 8:

[0540] Receiving and responding to feedback

[0541] The user previews the movie generated by the smartphone app and sends correction requests as necessary. Specifically, the user uses the feedback form on the preview screen to enter changes. The input data is the user's feedback. The output data is the feedback sent to the server as a correction request.

[0542] Step 9:

[0543] Final output and sharing

[0544] The server generates a final movie file incorporating the user's feedback and provides a download link. Specific operations include rendering the final movie file, generating a high-resolution file, and emailing the download link. The input data is the movie data with correction instructions. The output data is the final high-resolution movie file and a download link.

[0545] Through the above processing steps, users can easily create, save, and share moving and personalized movies based on digital data recordings of special moments.

[0546] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0547] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0548] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0549] [Second embodiment]

[0550] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0551] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0552] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0553] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0554] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0555] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0556] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0557] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0558] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0559] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0560] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0561] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0562] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, it is necessary to design and implement programs based on the following means.

[0563] 1. Data collection

[0564] Users upload photos, videos, and audio files to the system through dedicated applications or websites. The digital data provided is then sent to the server.

[0565] 2. Data Analysis and Classification

[0566] Server: Receives uploaded data and categorizes photos, videos, and audio recording files into their respective categories. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are converted to text using speech recognition technology to extract important messages and keywords.

[0567] 3. Build a story

[0568] Server: Based on the user's specified theme and wishes, the server automatically generates the optimal story based on key moments and messages. The AI ​​model builds the storyboard and determines the order and narrative of each scene.

[0569] 4. Music and Effects Selection

[0570] Server: Based on the story, select music from the library to enhance emotions and add effects appropriate for each scene. The music and effects are adjusted to fit the flow of the scene.

[0571] 5. Multilingual and culturally accommodating

[0572] Server: Performs multilingual text translation to generate subtitles and narration in the user's language, and customizes content based on the user's culture and customs.

[0573] 6. Privacy and Data Security

[0574] Server: Uploaded digital data is protected with the latest encryption technology, and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[0575] 7. Generate and preview the movie

[0576] Server: Integrates all elements and generates the final movie file, which is temporarily stored and made available for user preview.

[0577] User: Can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[0578] 8. Final output and sharing

[0579] Server: Generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g. high-resolution video file or streaming link) can be selected based on the user's preference.

[0580] Device: Receive the final movie file and save it locally or share it on an online platform.

[0581] Specific examples

[0582] For example, consider the case of creating a wedding movie.

[0583] 1. Data collection

[0584] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[0585] 2. Data Analysis and Classification

[0586] Server: Analyzes the bride and groom's appearance in photos and tags the moments of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes recorded conversations to extract moving messages.

[0587] 3. Build a story

[0588] Server: Arrange individual photos and video scenes on a storyboard to build a moving story from the bride and groom meeting to their wedding.

[0589] 4. Music and Effects Selection

[0590] Server: Insert moving music and apply effects (fade in / out, slow motion, etc.) to suit the scene.

[0591] 5. Multilingual and culturally accommodating

[0592] Server: For international weddings, generate subtitles and narration translated into each country's language.

[0593] 6. Privacy and Data Security

[0594] Server: All data is encrypted and secured.

[0595] 7. Generate and preview the movie

[0596] User: Preview the generated movie and request changes to the scene order or music.

[0597] 8. Final output and sharing

[0598] Server: Generates the final version of the movie and serves it to the user.

[0599] Device: Download movies, save them, and share them online.

[0600] In this way, the system according to the present invention automatically generates an inspiring movie based on the digital data recording a special moment of the user, and provides it as a personalized keepsake.

[0601] The processing flow will be explained below.

[0602] Step 1:

[0603] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The data is temporarily stored on the device.

[0604] Step 2:

[0605] Terminal: The data uploaded by the user is sent to the server via the Internet, including metadata (such as date, time, and location information).

[0606] Step 3:

[0607] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[0608] Step 4:

[0609] Server: Analyzes photo data using image recognition technology, automatically tags faces and distinctive scenes (e.g., smiles, touching moments).

[0610] Step 5:

[0611] Server: The video analysis module analyzes the video data, detects scene changes, and extracts important moments, such as the engagement kiss and ring exchange.

[0612] Step 6:

[0613] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[0614] Step 7:

[0615] Server: Generates a storyboard using key scenes and messages based on the user's specified theme and wishes. The AI ​​model determines the order of scenes and narrative progression to build a compelling story.

[0616] Step 8:

[0617] Server: Select music from the library to enhance the emotions of the story and add effects to each scene, including fade-in / out and slow motion.

[0618] Step 9:

[0619] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content for cultural adaptation.

[0620] Step 10:

[0621] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0622] Step 11:

[0623] Server: Integrates all elements and generates the initial movie version. The generated movie is temporarily saved and a preview link is provided to the user.

[0624] Step 12:

[0625] Users: Use the preview link to see the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[0626] Step 13:

[0627] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[0628] Step 14:

[0629] Server: Generates the final version of your movie and exports it in the format of your choice, including high-resolution video files and streaming links.

[0630] Step 15:

[0631] Terminal: Receive the final movie file and save it locally. If you want to share it on an online platform, you can easily do so using a dedicated application.

[0632] These are the basic processing steps in the embodiment of the present invention, which allows users to effectively record special moments and easily share them as moving movies.

[0633] Example 1

[0634] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0635] There is a need for a system that allows users to easily create inspiring and personalized movies. It is also necessary to meet the needs of each user for privacy protection, data security, multilingual support, and cultural adaptation. Conventional systems cannot comprehensively provide these features, and they are time-consuming and labor-intensive.

[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0637] In this invention, the server includes means for users to upload photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, means for automatically generating an inspiring story based on the user's wishes and themes, means for adding music and effects necessary for the inspiring story, means for providing multilingual support and cultural adaptation to the generated story, means for presenting the generated movie to the user and making corrections based on feedback, means for encrypting digital data to ensure privacy protection and data security, and means for exporting and sharing the final movie file in a specified format, thereby enabling users to easily create inspiring and personalized movies that also provide privacy protection, multilingual support, and cultural adaptation.

[0638] "Uploading" refers to the transfer of digital data from a user's terminal to a remote system such as a server.

[0639] "Data analysis" is the process of analyzing uploaded digital data, categorizing it into different formats such as photos, videos, and audio recording files, and understanding its content.

[0640] "Content classification" refers to the act of organizing analyzed digital data based on its type and characteristics, separating it into categories such as photos, videos, and audio recording files.

[0641] "Significant scene and message extraction" is the process of identifying and tagging particularly meaningful scenes and messages within photographs, videos, and audio recordings.

[0642] "Means for automated story generation" refers to algorithms and AI models that create compelling narrative structures from digital data based on user preferences and themes.

[0643] "Adding music and effects" is the process of applying emotionally enhancing music and visual effects to the generated story.

[0644] "Multilingual support" refers to the ability to translate generated content into multiple user-specified languages ​​and include subtitles and narration.

[0645] "Cultural adaptation" is the process of appropriately customizing content based on the user's specified culture and customs.

[0646] "Privacy protection" refers to various measures taken to protect users' personal information and digital data from others.

[0647] "Data security" means the technological measures and processes used to ensure the confidentiality, integrity and protection from unauthorized access of uploaded digital data.

[0648] "Export" refers to the act of converting the final generated movie file into a specific format (e.g., MP4, AVI, etc.) and outputting it.

[0649] "Sharing" refers to providing the generated movie file to other users or platforms so that they can access it.

[0650] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, the following hardware and software are used:

[0651] 1. Hardware and Software

[0652] Hardware: Servers, user devices (PCs, smartphones, etc.)

[0653] Software: Uploading applications and websites, image recognition technology (e.g., OpenCV, TensorFlow), speech recognition technology (e.g., Google Speech-to-Text API, IBM Watson), multilingual text translation systems (e.g., Google Translate API, DeepL), AI models (e.g., GPT-4)

[0654] 2. Detailed processing instructions

[0655] 1. Data collection

[0656] Users upload digital data (e.g., photos, videos, and audio files from the wedding day) through a dedicated application or website, and this data is sent to the server.

[0657] 2. Data Analysis and Classification

[0658] The server receives the uploaded data and categorizes it into photos, videos, and audio recordings. Image recognition technology is used to identify scenes in the photos and videos, confirm the bride and groom's appearance, and tag moments such as the vow kiss and ring exchange. Audio files are converted to text using speech recognition technology, which extracts important messages and keywords.

[0659] 3. Build a story

[0660] The server automatically generates the optimal story based on the user's specified theme and wishes, key moments, and messages. It uses a generative AI model (e.g., GPT-4) to build the storyboard and determine the order and narrative of each scene.

[0661] 4. Music and Effects Selection

[0662] The server selects inspiring music from a library and adds appropriate effects for each scene, with the music and effects tailored to the flow of the scene.

[0663] 5. Multilingual and culturally accommodating

[0664] The server uses a multilingual translation system to generate subtitles and narration in the language specified by the user, and also customizes the content based on the specified culture and customs.

[0665] 6. Privacy and Data Security

[0666] The server protects uploaded digital data using the latest encryption technology (e.g., AES-256), and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[0667] 7. Generate and preview the movie

[0668] The server aggregates all the elements and generates the final movie file using a video editing library (e.g., FFmpeg), which is temporarily stored in cloud storage and made available for users to preview.

[0669] 8. Final output and sharing

[0670] The server generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g., high-resolution video file or streaming link) can be selected based on the user's preference. The user's device receives the final movie file and can save it locally or share it on an online platform.

[0671] 3. Examples of concrete examples and prompts

[0672] A specific example of creating a wedding movie will be given below.

[0673] Data collection

[0674] Users upload photos and videos from their wedding day, as well as recorded messages from friends and family.

[0675] Data analysis and classification

[0676] The server analyzes the bride and groom's appearance in photos and tags the moment of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes audio messages to extract moving messages.

[0677] Building a story

[0678] The server uses these photos and video footage to build a moving story from the moment the bride and groom met to their wedding day.

[0679] Music and effects selection

[0680] Choose inspiring music and add effects that suit the scene (e.g. fade in / out, slow motion, etc.).

[0681] Multilingual and culturally accommodating

[0682] For international weddings, generate subtitles and narration translated into each country's language.

[0683] Privacy and Data Security

[0684] Encrypt your data to ensure the security of all your data.

[0685] Generate and preview the movie

[0686] The user can preview the generated movie and request changes to the scene order or music.

[0687] Final output and sharing

[0688] The server generates the final version of the movie and serves it to the user, who then downloads it, saves it, and shares it online.

[0689] Example prompt sentence:

[0690] "I would like to create a wedding video using the bride and groom's kiss and ring exchange scenes from the photos and videos I uploaded, and create a story with inspiring music. I would like the narration to be in both English and Spanish."

[0691] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0692] Step 1:

[0693] Data collection

[0694] Input: Photos, videos, and audio files from the user's wedding day

[0695] The user selects the digital data using a dedicated application or website and clicks an upload button.

[0696] The terminal transmits the selected data to a server via the Internet.

[0697] Output: Digital data sent to the server

[0698] Step 2:

[0699] Data analysis and classification

[0700] Input: Digital data sent to the server

[0701] The server runs a data analysis program that categorizes the photos, videos, and audio recordings into their respective categories.

[0702] The server uses image recognition technology (e.g., OpenCV, TensorFlow) to analyze the photos and video scenes and tag important moments such as the bride and groom's faces, the ring exchange, and the vow kiss.

[0703] The server uses voice recognition technology (e.g., Google Speech-to-Text API, IBM Watson) to convert the audio recording file into text and extract inspirational messages and keywords.

[0704] Output: Classified digital data, tagged key scenes, and audio data converted to text

[0705] Step 3:

[0706] Building a story

[0707] Input: Classified digital data, tagged important scenes, transcribed audio data, user preferences and themes

[0708] The server generates a storyboard using an AI model (e.g., GPT-4) based on the theme and wishes (prompt sentences) specified by the user.

[0709] The server places the tagged scenes and extracted messages on a storyboard and determines the order of each scene.

[0710] Output: Generated storyboard

[0711] Step 4:

[0712] Music and effects selection

[0713] Input: Generated storyboard

[0714] The server selects suitable songs from a music library (e.g., Epidemic Sound, Artlist) to enhance the emotion.

[0715] The server applies visual effects such as fade-in, fade-out, and slow motion to match the storyboard scenes.

[0716] Output: Story with added music and effects

[0717] Step 5:

[0718] Multilingual and culturally accommodating

[0719] Input: Stories with added music and effects, user-specified language and cultural preferences

[0720] The server uses a multilingual translation system (e.g., Google Translate API, DeepL) to translate subtitles and narration into the language specified by the user.

[0721] The server customizes the story content based on the specified culture and customs.

[0722] Output: Multilingual and culturally adapted stories

[0723] Step 6:

[0724] Privacy and Data Security

[0725] Input: Uploaded digital data

[0726] The server encrypts all digital data using the latest encryption technology (e.g., AES-256).

[0727] The server uses AI models to anonymize and mask unnecessary parts of the data to protect user privacy.

[0728] Output: Encrypted and privacy-protected digital data

[0729] Step 7:

[0730] Generate and preview the movie

[0731] Input: Multilingual and culturally adapted stories, encrypted and privacy-protected digital data

[0732] The server uses a video editing library (e.g. FFmpeg) to stitch all the elements together and generate the final movie file.

[0733] The server temporarily stores the generated movie file in cloud storage and provides a preview page URL for the user to preview it.

[0734] Users can visit a preview page to view the generated movie and provide feedback through the interface on any necessary corrections.

[0735] Output: Movie file with user review and feedback

[0736] Step 8:

[0737] Final output and sharing

[0738] Input: Movie file reviewed and feedback by user, user preferred format

[0739] The server makes corrections based on the user's feedback and generates the final version of the movie.

[0740] The server outputs the movie file in the user's desired export format (e.g., high-resolution video file, streaming link, etc.).

[0741] The user's device downloads the final movie file and stores it locally, allowing them to share it online via social media or email if desired.

[0742] Output: Final version movie file

[0743] (Application example 1)

[0744] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0745] This invention relates to a system that automatically generates moving, personalized movies simply by providing digital data (photos, videos, and audio recording files) from users. However, conventional systems lack specific application examples for further improving the user experience. In particular, they are unable to record shopping experiences in virtual stores and automatically generate movies from them, making it difficult for users to experience moving movies in different contexts. To solve this problem, it is necessary to expand the system to include movie generation based on shopping experiences in virtual stores.

[0746] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0747] In this invention, the server includes a means for uploading photos, videos, and audio recording files, a means for analyzing the uploaded data, categorizing the content, and extracting important scenes and messages, and a means for automatically generating an inspiring story based on the user's preferences and themes. This enables a shopping experience movie to be automatically generated using data taken by the user in a virtual store. The server also includes a means for performing facial recognition and tagging important moments, a means for tagging products and scenes in the virtual store, a means for converting conversation recording files into text using voice recognition technology and extracting important messages and keywords, and a means for generating inspiring prompts based on the user's shopping experience in the virtual store, further enriching the user's shopping experience and recording it as an inspiring movie.

[0748] "Photography" is a visual medium that uses light to record the shape and color of objects.

[0749] "Video" is a medium containing visual and audio information that records a time-sequence of images.

[0750] An "audio recording file" is a file that digitally records a person's voice or sound.

[0751] "Uploading" refers to the act of sending data owned by a user to a server via the Internet.

[0752] "Analysis" is the act of breaking down the provided data to understand its content and structure.

[0753] "Classification" is the act of dividing data into groups according to specific criteria.

[0754] An "important scene" is a moment that has particular meaning or impact in the overall flow or story.

[0755] A "message" is a portion of a voice recording that contains meaning or information extracted from the voice recording file.

[0756] "Automatic story generation" refers to the act of creating a coherent narrative based on analyzed data, based on the user's wishes and themes.

[0757] "Music" is sound with a melody or rhythm used to add emotional elements to a movie.

[0758] "Effects" are techniques and methods for giving a movie special visual or auditory effects.

[0759] "Multilingual support" refers to the ability to provide content in multiple languages.

[0760] "Cultural adaptation" is the act of adjusting content to suit a particular culture and customs.

[0761] "Privacy protection" refers to measures to protect personal information from unauthorized use.

[0762] "Data security" refers to the state in which data is protected from unauthorized access and tampering.

[0763] "Exporting" is the act of outputting the final movie file in a particular format.

[0764] "Sharing" is the act of making the generated movie available for viewing by others.

[0765] A "virtual store" is a virtual sales environment that offers products and services over the Internet.

[0766] "Shopping experience" is a general term for a series of actions and emotions that a user goes through when selecting and purchasing a product.

[0767] A "prompt sentence" is a sentence that is input to a generative AI model based on specified conditions and content.

[0768] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this invention, a system based on the following means is required.

[0769] 1. Data collection

[0770] The server provides a means for users to upload photos, videos, and audio recording files through dedicated applications or websites, making it easy for users to provide data.

[0771] 2. Data Analysis and Classification

[0772] The server receives the uploaded data and categorizes the photos, videos, and audio recordings into their respective categories. Image recognition technology is used to identify scenes in the photos and videos and tag specific moments. Audio recordings are converted into text using speech recognition technology, and important messages and keywords are extracted. Software such as OpenCV and TensorFlow are used for this process.

[0773] 3. Build a story

[0774] The server automatically generates the best story based on the user's specified theme and desires, key moments, and messages. It uses a generative AI model to build a storyboard and determine the order and narrative of each scene, creating a coherent and moving story.

[0775] 4. Music and Effects Selection

[0776] The server selects music from a library to enhance emotions based on the story, and adds effects appropriate for each scene. The music and effects are adjusted to match the flow of the scenes. This process is performed using software such as PIL (Python Imaging Library) and moviepy.

[0777] 5. Multilingual and culturally accommodating

[0778] The server performs multilingual text translation to generate subtitles and narration in the user's language, and customizes the content based on the user's culture and customs. By using translation software such as Google Translate API, it can accommodate users from different cultures.

[0779] 6. Privacy and Data Security

[0780] The server protects uploaded digital data using the latest encryption technology. AI models also anonymize data and mask unnecessary parts to maximize user privacy. Security software such as AES (Advanced Encryption Standard) is used.

[0781] 7. Generate and preview the movie

[0782] The server combines all the elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. Users can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[0783] 8. Final output and sharing

[0784] The server generates the final version of the movie incorporating the user's feedback and corrections. The user can choose the export format of the movie file (e.g., high-resolution video file or streaming link) according to their preference. The user receives the final movie file and can save it locally or share it on an online platform.

[0785] As a concrete example, consider a scenario where a user uploads photos and videos of clothes they tried on in a virtual store, and a shopping experience video is generated based on those photos and videos. Here is an example of a prompt to be input to the generative AI model:

[0786] "Create an inspiring shopping experience video using photos and videos of users trying on different outfits in a virtual store. At the end of the video, highlight the user's favorite outfit and play inspiring music in the background."

[0787] This prompt allows the generative AI model to generate a user experience movie that can then be previewed and shared.

[0788] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0789] Step 1:

[0790] Users upload photos, videos, and audio files through dedicated applications or websites. The input is the user's digital data, and the output is data stored on the server. Specifically, the user selects the captured data and presses the upload button, which sends the file to the server.

[0791] Step 2:

[0792] The server analyzes the uploaded data and classifies photos, videos, and audio recording files into their respective categories. The input is the uploaded digital data, and the output is the classified data. Specifically, image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are also converted into text using speech recognition technology, and important messages and keywords are extracted. This processing is done using OpenCV and TensorFlow.

[0793] Step 3:

[0794] The server references the user's specified theme and wishes, and automatically generates an optimal story based on key moments and messages. The input is classified data and the user's theme settings, and the output is a storyboard. Specifically, a generative AI model analyzes the data and constructs the story flow. This process generates a coherent and moving story.

[0795] Step 4:

[0796] Based on the story, the server selects music from a library to enhance emotions and adds appropriate effects to each scene. The input is a storyboard, and the output is the final video scene. Specifically, PIL (Python Imaging Library) and moviepy are used to insert music and effects into the video. This process creates visually and aurally appealing content.

[0797] Step 5:

[0798] The server performs multilingual text translation to generate subtitles and narration in the user's specified language. It also customizes the content based on the specified culture and customs. The input is the story text information and the user's specified language, and the output is the translated subtitles and narration. Specifically, it uses the Google Translate API or similar to translate the text data and generate the required subtitles and narration.

[0799] Step 6:

[0800] The server protects uploaded digital data using the latest encryption technology. Furthermore, the AI ​​model anonymizes the data and masks unnecessary parts to ensure maximum user privacy. The input is the user's digital data, and the output is protected data. Specifically, the data is encrypted using security technologies such as AES (Advanced Encryption Standard).

[0801] Step 7:

[0802] The server aggregates all elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. The input is the aggregated content, and the output is the generated movie file. Specifically, software such as moviepy is used to combine all scenes and effects and create the final movie file.

[0803] Step 8:

[0804] The user previews the generated movie and provides feedback through the interface on any necessary modifications (e.g., changing the order of scenes or replacing music). The input is the generated movie file and the user's feedback, and the output is the modified movie file. Specifically, the server performs the modifications based on the feedback provided by the user.

[0805] Step 9:

[0806] The server generates the final version of the movie incorporating the modifications based on the user's feedback. The export format (e.g., high-resolution video file or streaming link) can be selected according to the user's preference. The input is the modified movie file, and the output is the final movie file. Specifically, we provide the function to export the movie file in the format desired by the user.

[0807] Step 10:

[0808] Users receive the final movie file and can save it locally or share it on an online platform. The input is the final movie file, and the output is the shared movie. Specifically, the generated movie file can be uploaded to social media or cloud storage to be shared with other users.

[0809] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0810] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by the user. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more emotionally rich stories. To implement the system, it is necessary to design and implement programs based on the following means.

[0811] 1. Data collection

[0812] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. These data are temporarily stored on the device.

[0813] 2. Data transmission

[0814] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[0815] 3. Data Analysis and Classification

[0816] Server: Receives uploaded data and classifies photos, videos, and audio recordings. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recordings are also converted to text using speech recognition technology to extract important messages and keywords.

[0817] 4. Emotional Recognition

[0818] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[0819] 5. Build a story

[0820] Server: Based on the user's specified themes and wishes, and taking into account the emotional information recognized by the emotion engine, the server automatically generates a moving story. The AI ​​model determines the appropriate scene order and narrative to build a compelling story.

[0821] 6. Music and Effects Selection

[0822] Server: Based on the story, select music from the library to enhance the emotion and add effects to suit the scene, such as fade in / out and slow motion.

[0823] 7. Multilingual and culturally adaptable

[0824] Server: Generates subtitles and narration based on the user-specified language. Uses text translation tools for multilingual support and adjusts content for cultural adaptation.

[0825] 8. Privacy and Data Security

[0826] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0827] 9. Generate and preview the movie

[0828] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and users can view it using the preview link.

[0829] Users: can use the preview link to see the generated movie and provide feedback, such as requests for changes to the scene order or music.

[0830] 10. Receiving and Responding to Feedback

[0831] Server: Receives user feedback and modifies the movie based on the instructions. The modified movie is then regenerated and saved.

[0832] 11. Final output and sharing

[0833] Server: Generates the final version of the movie and provides it to the user in the export format of their choice, such as a high-resolution video file or a streaming link.

[0834] Device: Receive the final movie file and save it locally or share it on an online platform.

[0835] Specific examples

[0836] For example, consider the case of creating a wedding movie.

[0837] 1. Data collection

[0838] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[0839] 2. Data transmission

[0840] Terminal: Sends these uploaded data to the server.

[0841] 3. Data Analysis and Classification

[0842] Server: Recognizes and tags important moments in photos and videos (e.g., the promise kiss, the ring exchange), and extracts inspirational messages from audio recordings.

[0843] 4. Emotional Recognition

[0844] Server: Analyzes facial expressions and tone of voice from video and audio to recognize the emotions of the bride and groom and guests. For example, it identifies touching moments by detecting smiles and tears.

[0845] 5. Build a story

[0846] Server: Based on the extracted scenes and messages and the emotional information recognized by the emotion engine, a moving story of the entire wedding is constructed.

[0847] 6. Music and Effects Selection

[0848] Server: Add inspiring music and apply effects appropriate to the scene, including fade in / out, slow motion, etc.

[0849] 7. Multilingual and culturally adaptable

[0850] Server: For international weddings, generate subtitles and narration translated into each country's language.

[0851] 8. Privacy and Data Security

[0852] Server: All data is encrypted and secured.

[0853] 9. Generate and preview the movie

[0854] User: Preview the generated movie and request corrections if necessary.

[0855] 10. Receiving and Responding to Feedback

[0856] Server: Based on user feedback, the movie is regenerated and modifications are made.

[0857] 11. Final output and sharing

[0858] Server: Generates the final version of the movie and serves it to the user.

[0859] Device: Download movies and finally save and share them.

[0860] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[0861] The processing flow will be explained below.

[0862] Step 1:

[0863] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. Uploaded data is temporarily stored on the device.

[0864] Step 2:

[0865] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[0866] Step 3:

[0867] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[0868] Step 4:

[0869] Server: Analyzes photo data using image recognition technology, automatically identifying faces and tagging distinctive scenes (e.g., smiles or touching moments).

[0870] Step 5:

[0871] Server: Analyzes video data using the video analysis module, detects scene changes, and extracts important moments (e.g., the engagement kiss or ring exchange).

[0872] Step 6:

[0873] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[0874] Step 7:

[0875] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[0876] Step 8:

[0877] Server: Automatically generates an emotional storyboard based on the user's specified themes and wishes, and also takes into account recognized emotional information. The AI ​​model determines the order of scenes and the progression of the story, building a coherent story.

[0878] Step 9:

[0879] Server: Based on the story, select music from the library to enhance the emotions and add effects to each scene, such as fade in / out and slow motion.

[0880] Step 10:

[0881] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content to take cultural adaptation into account.

[0882] Step 11:

[0883] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[0884] Step 12:

[0885] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and a preview link is provided to users.

[0886] Step 13:

[0887] Users: Use the preview link to view the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[0888] Step 14:

[0889] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[0890] Step 15:

[0891] Server: Generates the final movie file and exports it in the format of your choice (e.g. high-resolution video file or streaming link).

[0892] Step 16:

[0893] Terminal: Receive the final movie file and save it locally. Furthermore, if you want to share it on an online platform, you can easily do so using a dedicated application.

[0894] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[0895] Example 2

[0896] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0897] In recent years, the widespread use of smartphones and digital cameras has increased the opportunities for users to save various events and everyday moments as digital data (photos, videos, and audio recording files). However, editing this data to create moving and personalized movies requires advanced editing skills and time, making it difficult for average users. Conventional methods have been problematic in that the editing process is cumbersome and it is difficult to automatically generate emotionally appealing stories, leaving users unable to obtain satisfactory results.

[0898] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to upload photos, videos, and audio recording files; a means for transmitting the uploaded data to the server via the Internet; a means for the server to analyze the received data, classify the content, and extract important scenes and messages; a means for analyzing the user's emotions using emotion recognition technology; a means for automatically generating a story based on the user's preferences and themes, taking into account the emotional information; a means for adding music and effects to the story; a means for multilingual support and cultural adaptation; a means for presenting the generated movie to the user and making corrections based on the user's feedback; a means for ensuring privacy protection and data security; and a means for exporting and sharing the final movie file in a specified format. This allows even ordinary users to easily automatically generate moving and personalized movies and obtain high-quality results.

[0899] "User" refers to any individual or organization that uses this system.

[0900] "Photo" refers to still image data that a user takes using a digital device and uploads to the system.

[0901] "Video" refers to moving image data that a user shoots using a digital device and uploads to the system.

[0902] "Audio recording file" refers to a file that contains audio data recorded by a user or another person.

[0903] "Upload" refers to a user sending a photo, video, or audio recording file to the system through a dedicated application or website.

[0904] "Internet" refers to a global network for data communication.

[0905] "Server" refers to a computer system that receives, analyzes, sorts, and processes data sent by users.

[0906] "Analysis" refers to the process by which the server understands and interprets the content of photographs, videos, and audio recordings based on data.

[0907] "Classification" refers to the process by which the server divides the data it receives into specific categories.

[0908] "Important scenes and messages" refer to notable moments or words that the user finds interesting or that the system automatically selects.

[0909] "Emotion recognition technology" refers to the technology that a system uses to analyze and identify a user's emotions from images and audio.

[0910] "Automatic story generation" refers to the process in which a system constructs a series of events into a narrative based on the user's wishes, themes, and emotional information.

[0911] "Music and Effects" refers to background music and visual effects added to enhance the visual and auditory impression of the movie.

[0912] "Multilingual" refers to the ability to display content or provide narration in different languages.

[0913] "Cultural adaptation" refers to adjusting content to suit the user's culture.

[0914] "Feedback" refers to the act of a user conveying to the system their opinions and requests for corrections and improvements to a movie.

[0915] "Privacy protection" refers to protecting users' personal information and data from being leaked to third parties.

[0916] "Data security" refers to measures taken to keep data protected from unauthorized access and tampering.

[0917] "Export" refers to the act of saving or outputting the final movie file in a particular format.

[0918] "Sharing" refers to sharing the generated movie file or link with other users and platforms.

[0919] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. Specific procedures for implementing the present invention and the hardware and software used are described below.

[0920] This system is primarily composed of users, devices, and a server. Users upload data using a dedicated application or website, and their devices send the data to the server via the Internet. The server analyzes the data, generates a story, and provides the final movie file to the user. The specific operation is explained below.

[0921] Data collection

[0922] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website, and these data are temporarily stored on the device.

[0923] Sending data

[0924] The device sends the data uploaded by the user to a server via the Internet, including metadata (e.g., date, time, location, etc.).

[0925] Data analysis and classification

[0926] The server analyzes the received data and classifies each photo, video, and audio recording file using image recognition technologies (e.g., Google Cloud Vision API) and speech recognition technologies (e.g., Google Cloud Speech-to-Text). Scenes in the photos and videos are identified and specific moments (e.g., the engagement kiss or the ring exchange) are tagged. Audio recordings are converted to text and key messages and keywords are extracted.

[0927] Emotion recognition

[0928] The server uses emotion recognition technology (e.g., Microsoft Azure Emotion API) to recognize the user's emotions from video and audio, analyzing smiles, tears, tone of voice, etc. to identify emotional moments.

[0929] Building a story

[0930] The server automatically generates a moving story based on the user's preferences and themes, taking into account emotional information. This story generation uses an AI model (e.g., OpenAI GPT) to determine the order of movie scenes and the content of the narration.

[0931] Music and effects selection

[0932] Based on the story, the server selects music from a library (e.g., Epidemic Sound) to enhance the emotions and adds effects (fade in / out, slow motion, etc.) that match the scene.

[0933] Multilingual and culturally accommodating

[0934] The server generates subtitles and narration based on the language specified by the user, and uses text translation tools (e.g., Google Translate API) to support multiple languages, adjusting the content for cultural adaptation.

[0935] Privacy and Data Security

[0936] The server uses the latest encryption technology (e.g., AES-256) for all uploaded data to ensure privacy and data security. Data is sent and received using SSL / TLS encryption, and data is encrypted when stored.

[0937] Generate and preview the movie

[0938] The server combines all the elements to generate a first version of the movie file, stores it temporarily, and makes it available to users via a preview link.

[0939] Users can use the preview link to view the resulting movie and submit requests for changes to the scene order or music, if necessary.

[0940] Receiving and responding to feedback

[0941] The server receives feedback from the user and modifies the movie based on the instructions, and the modified movie is regenerated and saved.

[0942] Final output and sharing

[0943] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link.

[0944] The device will receive the final movie file and store it locally, or you can share it on an online platform.

[0945] Examples of prompt statements

[0946] Below are some example prompts to input to the generative AI model:

[0947] Create a moving wedding video using photos and videos from your wedding, along with messages from friends and family. Highlight the bride and groom's smiling and tearful moments, and add music appropriate to their special day. For an international wedding, add subtitles in English and Japanese to reflect cultural elements. Specifically, use emotional music for the groom's entrance and select a song that gradually builds up during the bride's entrance. Additionally, highlight the vows and ring exchange in slow motion, and add moving guest messages.

[0948] Through the above steps, the present invention utilizes an emotion engine to automatically generate an even more moving movie based on digital data that records a user's special moments, and provides it as a personalized keepsake.

[0949] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0950] Step 1: Collect data

[0951] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The input is the user's own digital data. The uploaded data is temporarily stored on the device and becomes the output. For example, a user can drag and drop photos from their wedding day or audio files containing messages from friends into the application, and then enter the event theme and desired details into a form.

[0952] Step 2: Sending data

[0953] The device sends the data uploaded by the user to the server via the Internet. The input includes data temporarily stored on the device and metadata (date, time, location, etc.). The output is data sent to the server and received by the server. Specifically, the device securely sends this data to the server using the HTTPS protocol.

[0954] Step 3: Analyze and classify the data

[0955] The server analyzes and classifies the data it receives. The input is the data received by the server. The output is photos, videos, and audio recordings sorted into their respective categories, with each file tagged with important scenes and keywords. For example, the server uses Google Cloud Vision API to identify scenes in photos and videos and tag specific moments, such as the bride and groom's kiss and ring exchange. It also uses Google Cloud Speech-to-Text to convert audio recordings into text and extract important messages and keywords.

[0956] Step 4: Recognize emotions

[0957] The server uses emotion recognition technology to analyze user emotions from video and audio. The input is analyzed and classified video and audio data. The output is an emotional tag attached to each piece of data. Specifically, it uses the Microsoft Azure Emotion API to analyze smiles, tears, tone of voice, and other emotions to identify moments when the bride and groom and guests are emotional.

[0958] Step 5: Build your story

[0959] The server automatically generates a story based on the user's wishes and themes, taking into account emotional information. The inputs are data tagged with emotions and the wishes and themes provided by the user. The output is a movie plan that determines the order of scenes and narration for a moving story. Specifically, it uses OpenAI GPT to construct the story's narrative and determine the placement of scenes.

[0960] Step 6: Choose music and effects

[0961] The server selects music and effects based on the story. The input is a story plan. The output is a story plan with music and effects applied to scenes to enhance the emotions. Specific operations include selecting appropriate music from the Epidemic Sound library and applying effects such as fade-in / out and slow motion to scenes.

[0962] Step 7: Multilingualism and cultural adaptation

[0963] The server generates subtitles and narration based on the language specified by the user. The inputs are the language information specified by the user and a story plan. The output is a story plan that includes subtitles and narration in multiple languages. Specifically, the system uses the Google Translate API to translate the text into multiple languages ​​and edits the content to take cultural adaptation into account.

[0964] Step 8: Privacy and Data Security

[0965] The server uses encryption technology when sending, receiving, and storing data. The inputs include user data stored on the server and data being sent and received. The output is encrypted data. Specifically, data security is ensured using AES-256 data encryption and SSL / TLS protocols.

[0966] Step 9: Generate and preview your movie

[0967] The server integrates all elements to generate a first-run movie file and temporarily stores it. The input is the final story plan. The output is the first-run movie file. Specifically, video editing software is used to generate the movie based on the pre-planned layout.

[0968] The user uses the preview link to view the generated movie. The input is the generated movie file. The output is feedback based on the preview results.

[0969] Step 10: Receiving and responding to feedback

[0970] The server receives user feedback and modifies the movie based on the instructions. The inputs are the user's feedback and the original movie file. The output is a modified movie file that reflects the feedback. Specifically, the server reorders the requested scenes and adjusts the music to generate a new movie.

[0971] Step 11: Final output and sharing

[0972] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link. The input is the final, modified movie plan. The output is the final movie file. Specific operations include generating an HD MP4 file or a YouTube streaming link as the final output.

[0973] The device receives the final movie file and stores it locally. The input is the final movie file provided by the server. The output is a locally stored movie file, which the user can also upload to an online platform for sharing.

[0974] (Application example 2)

[0975] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0976] Conventional digital content generation systems that use photos, videos, and audio files have struggled to automatically generate emotionally rich stories. Furthermore, there are challenges, such as generating movies that effectively reflect the user's emotions, supporting multiple languages, and protecting privacy. In particular, there is a need for a system that can visually and emotionally recreate users' memories and share them in high resolution with simple smartphone application operations.

[0977] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, and means for automatically generating an emotional story based on the user's wishes and themes. This makes it possible to automatically generate an emotional digital story that effectively reflects the user's emotions using an emotion engine.

[0978] A "photograph" is a file that digitally records a still image.

[0979] A "video" is a file that contains a digital recording of moving images.

[0980] An "audio recording file" is a file that records audio in digital format.

[0981] "Uploading" means sending local data owned by the user to a server via the Internet.

[0982] "Data analysis" refers to the use of machine learning and algorithms to process uploaded digital content and extract specific information.

[0983] "Content classification" refers to grouping photos, videos, and audio recording files based on their format and characteristics.

[0984] "Key Scene and Message Extraction" refers to automatically identifying moments and meaningful messages within digital content that deserve special emphasis.

[0985] "User's wishes and themes" are requests and objectives for a story or movie specified by the user.

[0986] "Automatic generation of emotional stories" is the process of automatically constructing emotional stories based on extracted scenes and messages.

[0987] An "emotion engine" is an algorithm that recognizes user emotions from digital data and uses that information to create stories and movies.

[0988] "Adding Music and Effects" is the process of adding music and visual effects to a moving story to enhance its emotional expression.

[0989] "Multilingual support" means translating generated content into multiple languages ​​to accommodate users from different cultures and regions.

[0990] "Cultural adaptation" means adjusting content to suit users in a particular culture or region.

[0991] "Privacy protection" refers to measures and technologies to protect user data and information from third parties.

[0992] "Data security" means ensuring that digital data is protected and that unauthorized access is prevented.

[0993] "Final Movie File" is the finished digital movie file exported in the user's desired format.

[0994] "Export in specified format" means outputting the final movie in a file format selected by the user.

[0995] "Emotion analysis" refers to recognizing and identifying a user's emotional state from digital data.

[0996] "Previewing a movie" means checking a movie before it is completed.

[0997] "Accepting a change request" means receiving a request for correction or addition from a user and changing the content.

[0998] A "high resolution movie file" is a digital movie file with high quality video output.

[0999] "Share on social media" means posting and sharing the generated movie file on an internet communication platform.

[1000] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. By combining this system with an emotion engine, it creates a story that reflects the user's emotions.

[1001] Program Generation and Processing Description

[1002] Hardware and software used:

[1003] Server: High performance data processing and storage server.

[1004] Terminal (smartphone or tablet): A device through which a user uploads digital data.

[1005] Network Connection: An internet connection for uploading, downloading, and communicating data.

[1006] Software used:

[1007] TensorFlow: A machine learning framework for image and speech recognition.

[1008] OpenCV: A library for image and video analysis.

[1009] Django: A server-side web framework.

[1010] React Native: A development platform for smartphone applications.

[1011] Amazon Web Services (AWS) S3: Data storage service.

[1012] Amazon Polly: A speech synthesis service.

[1013] Examples of implementation:

[1014] 1. Uploading data:

[1015] Users use a smartphone app to upload photos, videos, and audio recordings related to events such as weddings and birthdays, and metadata (date, time, and location information) is automatically captured and added to the uploaded data.

[1016] 2. Data transmission:

[1017] The device sends the data uploaded by the user to a server via the Internet, where it is stored in Amazon Web Services (AWS) S3.

[1018] 3. Data analysis and classification:

[1019] The server receives uploaded data using the Django framework and leverages TensorFlow and OpenCV to analyze and classify photo, video, and audio recording files, automatically extracting and tagging important scenes and messages from the uploaded data.

[1020] 4. Emotion Recognition:

[1021] The server uses an emotion engine to analyze emotions in the data, for example by analyzing facial expressions and tone of voice in the video to identify moments of emotion or joy.

[1022] 5. Build a story:

[1023] The server automatically generates an emotional story based on the user's specified theme and wishes, and this process incorporates the emotional information recognized by the emotion engine.

[1024] 6. Add music and effects:

[1025] The server adds appropriate music and effects to the generated story, and also generates voice narration using Amazon Polly, including multilingual support and cultural adaptation.

[1026] 7. Generate and preview the movie:

[1027] The server generates the initial movie file, allows users to preview it through a smartphone app, and accepts and responds to user requests for changes.

[1028] 8. Final movie output:

[1029] Users can download the final version of their movie in high resolution and share it on social media, all while ensuring privacy and data security.

[1030] Examples of prompts:

[1031] "Analyze family trip data to create an inspiring movie that captures moments of happiness for users."

[1032] In this way, the system of the present invention utilizes digital data of special moments provided by the user to automatically generate moving and personalized movies, allowing memories to be recreated in a deeper and more emotional way.

[1033] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1034] Step 1:

[1035] Data collection and upload

[1036] Users use a smartphone application to upload photos, videos, and audio recordings. Specifically, they tap the app's "upload" button and select media files from their device. The input data are photos, videos, and audio recordings and their associated metadata. The output data is digital data sent to a server.

[1037] Step 2:

[1038] Sending data

[1039] The device sends the data uploaded by the user to a server via the Internet. Specifically, the application transfers the selected file to the server using the HTTPS protocol. The input data is the digital content specified by the user through the application. The output data is the file stored on the server.

[1040] Step 3:

[1041] Data analysis and classification

[1042] The server receives, analyzes, and classifies the uploaded data. It uses TensorFlow and OpenCV to analyze photo, video, and audio recording files and processes the data according to their respective formats. Specific operations include scene analysis using image recognition and text conversion using voice recognition. The input data is the uploaded digital data. The output data is the classified content and its metadata.

[1043] Step 4:

[1044] Emotion recognition

[1045] The server uses an emotion engine to analyze users' emotions from uploaded videos and audio. Specifically, it analyzes facial expressions and tone of voice to identify emotional elements and emotional moments. The input data are classified photos, videos, and audio recordings. The output data is media data with recognized emotion information.

[1046] Step 5:

[1047] Building a story

[1048] The server automatically generates an inspiring story based on the user's specified theme and wishes. It primarily uses an AI model to construct the story based on recognized emotional information and tagged scenes. Specifically, it determines the order of scenes and automatically forms a narrative. The input data is media data with emotional information and the user's specified theme. The output data is the constructed story.

[1049] Step 6:

[1050] Adding music and effects

[1051] The server adds appropriate music and visual effects to the generated story. It uses Amazon Polly to generate narration, and also supports multiple languages ​​and cultural adaptations. Specific operations include selecting music, applying effects, and inserting narration. The input data is the constructed story. The output data is a movie with music and effects added.

[1052] Step 7:

[1053] Generate and preview the movie

[1054] The server generates the final movie file and allows users to preview it through a smartphone app. Specific operations include rendering the movie file and generating a preview file. The input data is movie data with music and effects added. The output data is a movie file that users can preview.

[1055] Step 8:

[1056] Receiving and responding to feedback

[1057] The user previews the movie generated by the smartphone app and sends correction requests as necessary. Specifically, the user uses the feedback form on the preview screen to enter changes. The input data is the user's feedback. The output data is the feedback sent to the server as a correction request.

[1058] Step 9:

[1059] Final output and sharing

[1060] The server generates a final movie file incorporating the user's feedback and provides a download link. Specific operations include rendering the final movie file, generating a high-resolution file, and emailing the download link. The input data is the movie data with correction instructions. The output data is the final high-resolution movie file and a download link.

[1061] Through the above processing steps, users can easily create, save, and share moving and personalized movies based on digital data recordings of special moments.

[1062] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1063] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1064] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1065] [Third embodiment]

[1066] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1067] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1068] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1069] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1070] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1071] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1072] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1073] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1074] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1075] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1076] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1077] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1078] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, it is necessary to design and implement programs based on the following means.

[1079] 1. Data collection

[1080] Users upload photos, videos, and audio files to the system through dedicated applications or websites. The digital data provided is then sent to the server.

[1081] 2. Data Analysis and Classification

[1082] Server: Receives uploaded data and categorizes photos, videos, and audio recording files into their respective categories. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are converted to text using speech recognition technology to extract important messages and keywords.

[1083] 3. Build a story

[1084] Server: Based on the user's specified theme and wishes, the server automatically generates the optimal story based on key moments and messages. The AI ​​model builds the storyboard and determines the order and narrative of each scene.

[1085] 4. Music and Effects Selection

[1086] Server: Based on the story, select music from the library to enhance emotions and add effects appropriate for each scene. The music and effects are adjusted to fit the flow of the scene.

[1087] 5. Multilingual and culturally accommodating

[1088] Server: Performs multilingual text translation to generate subtitles and narration in the user's language, and customizes content based on the user's culture and customs.

[1089] 6. Privacy and Data Security

[1090] Server: Uploaded digital data is protected with the latest encryption technology, and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[1091] 7. Generate and preview the movie

[1092] Server: Integrates all elements and generates the final movie file, which is temporarily stored and made available for user preview.

[1093] User: Can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[1094] 8. Final output and sharing

[1095] Server: Generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g. high-resolution video file or streaming link) can be selected based on the user's preference.

[1096] Device: Receive the final movie file and save it locally or share it on an online platform.

[1097] Specific examples

[1098] For example, consider the case of creating a wedding movie.

[1099] 1. Data collection

[1100] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[1101] 2. Data Analysis and Classification

[1102] Server: Analyzes the bride and groom's appearance in photos and tags the moments of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes recorded conversations to extract moving messages.

[1103] 3. Build a story

[1104] Server: Arrange individual photos and video scenes on a storyboard to build a moving story from the bride and groom meeting to their wedding.

[1105] 4. Music and Effects Selection

[1106] Server: Insert moving music and apply effects (fade in / out, slow motion, etc.) to suit the scene.

[1107] 5. Multilingual and culturally accommodating

[1108] Server: For international weddings, generate subtitles and narration translated into each country's language.

[1109] 6. Privacy and Data Security

[1110] Server: All data is encrypted and secured.

[1111] 7. Generate and preview the movie

[1112] User: Preview the generated movie and request changes to the scene order or music.

[1113] 8. Final output and sharing

[1114] Server: Generates the final version of the movie and serves it to the user.

[1115] Device: Download movies, save them, and share them online.

[1116] In this way, the system according to the present invention automatically generates an inspiring movie based on the digital data recording a special moment of the user, and provides it as a personalized keepsake.

[1117] The processing flow will be explained below.

[1118] Step 1:

[1119] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The data is temporarily stored on the device.

[1120] Step 2:

[1121] Terminal: The data uploaded by the user is sent to the server via the Internet, including metadata (such as date, time, and location information).

[1122] Step 3:

[1123] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[1124] Step 4:

[1125] Server: Analyzes photo data using image recognition technology, automatically tags faces and distinctive scenes (e.g., smiles, touching moments).

[1126] Step 5:

[1127] Server: The video analysis module analyzes the video data, detects scene changes, and extracts important moments, such as the engagement kiss and ring exchange.

[1128] Step 6:

[1129] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[1130] Step 7:

[1131] Server: Generates a storyboard using key scenes and messages based on the user's specified theme and wishes. The AI ​​model determines the order of scenes and narrative progression to build a compelling story.

[1132] Step 8:

[1133] Server: Select music from the library to enhance the emotions of the story and add effects to each scene, including fade-in / out and slow motion.

[1134] Step 9:

[1135] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content for cultural adaptation.

[1136] Step 10:

[1137] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1138] Step 11:

[1139] Server: Integrates all elements and generates the initial movie version. The generated movie is temporarily saved and a preview link is provided to the user.

[1140] Step 12:

[1141] Users: Use the preview link to see the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[1142] Step 13:

[1143] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[1144] Step 14:

[1145] Server: Generates the final version of your movie and exports it in the format of your choice, including high-resolution video files and streaming links.

[1146] Step 15:

[1147] Terminal: Receive the final movie file and save it locally. If you want to share it on an online platform, you can easily do so using a dedicated application.

[1148] These are the basic processing steps in the embodiment of the present invention, which allows users to effectively record special moments and easily share them as moving movies.

[1149] Example 1

[1150] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1151] There is a need for a system that allows users to easily create inspiring and personalized movies. It is also necessary to meet the needs of each user for privacy protection, data security, multilingual support, and cultural adaptation. Conventional systems cannot comprehensively provide these features, and they are time-consuming and labor-intensive.

[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1153] In this invention, the server includes means for users to upload photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, means for automatically generating an inspiring story based on the user's wishes and themes, means for adding music and effects necessary for the inspiring story, means for providing multilingual support and cultural adaptation to the generated story, means for presenting the generated movie to the user and making corrections based on feedback, means for encrypting digital data to ensure privacy protection and data security, and means for exporting and sharing the final movie file in a specified format, thereby enabling users to easily create inspiring and personalized movies that also provide privacy protection, multilingual support, and cultural adaptation.

[1154] "Uploading" refers to the transfer of digital data from a user's terminal to a remote system such as a server.

[1155] "Data analysis" is the process of analyzing uploaded digital data, categorizing it into different formats such as photos, videos, and audio recording files, and understanding its content.

[1156] "Content classification" refers to the act of organizing analyzed digital data based on its type and characteristics, separating it into categories such as photos, videos, and audio recording files.

[1157] "Significant scene and message extraction" is the process of identifying and tagging particularly meaningful scenes and messages within photographs, videos, and audio recordings.

[1158] "Means for automated story generation" refers to algorithms and AI models that create compelling narrative structures from digital data based on user preferences and themes.

[1159] "Adding music and effects" is the process of applying emotionally enhancing music and visual effects to the generated story.

[1160] "Multilingual support" refers to the ability to translate generated content into multiple user-specified languages ​​and include subtitles and narration.

[1161] "Cultural adaptation" is the process of appropriately customizing content based on the user's specified culture and customs.

[1162] "Privacy protection" refers to various measures taken to protect users' personal information and digital data from others.

[1163] "Data security" means the technological measures and processes used to ensure the confidentiality, integrity and protection from unauthorized access of uploaded digital data.

[1164] "Export" refers to the act of converting the final generated movie file into a specific format (e.g., MP4, AVI, etc.) and outputting it.

[1165] "Sharing" refers to providing the generated movie file to other users or platforms so that they can access it.

[1166] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, the following hardware and software are used:

[1167] 1. Hardware and Software

[1168] Hardware: Servers, user devices (PCs, smartphones, etc.)

[1169] Software: Uploading applications and websites, image recognition technology (e.g., OpenCV, TensorFlow), speech recognition technology (e.g., Google Speech-to-Text API, IBM Watson), multilingual text translation systems (e.g., Google Translate API, DeepL), AI models (e.g., GPT-4)

[1170] 2. Detailed processing instructions

[1171] 1. Data collection

[1172] Users upload digital data (e.g., photos, videos, and audio files from the wedding day) through a dedicated application or website, and this data is sent to the server.

[1173] 2. Data Analysis and Classification

[1174] The server receives the uploaded data and categorizes it into photos, videos, and audio recordings. Image recognition technology is used to identify scenes in the photos and videos, confirm the bride and groom's appearance, and tag moments such as the vow kiss and ring exchange. Audio files are converted to text using speech recognition technology, which extracts important messages and keywords.

[1175] 3. Build a story

[1176] The server automatically generates the optimal story based on the user's specified theme and wishes, key moments, and messages. It uses a generative AI model (e.g., GPT-4) to build the storyboard and determine the order and narrative of each scene.

[1177] 4. Music and Effects Selection

[1178] The server selects inspiring music from a library and adds appropriate effects for each scene, with the music and effects tailored to the flow of the scene.

[1179] 5. Multilingual and culturally accommodating

[1180] The server uses a multilingual translation system to generate subtitles and narration in the language specified by the user, and also customizes the content based on the specified culture and customs.

[1181] 6. Privacy and Data Security

[1182] The server protects uploaded digital data using the latest encryption technology (e.g., AES-256), and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[1183] 7. Generate and preview the movie

[1184] The server aggregates all the elements and generates the final movie file using a video editing library (e.g., FFmpeg), which is temporarily stored in cloud storage and made available for users to preview.

[1185] 8. Final output and sharing

[1186] The server generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g., high-resolution video file or streaming link) can be selected based on the user's preference. The user's device receives the final movie file and can save it locally or share it on an online platform.

[1187] 3. Examples of concrete examples and prompts

[1188] A specific example of creating a wedding movie will be given below.

[1189] Data collection

[1190] Users upload photos and videos from their wedding day, as well as recorded messages from friends and family.

[1191] Data analysis and classification

[1192] The server analyzes the bride and groom's appearance in photos and tags the moment of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes audio messages to extract moving messages.

[1193] Building a story

[1194] The server uses these photos and video footage to build a moving story from the moment the bride and groom met to their wedding day.

[1195] Music and effects selection

[1196] Choose inspiring music and add effects that suit the scene (e.g. fade in / out, slow motion, etc.).

[1197] Multilingual and culturally accommodating

[1198] For international weddings, generate subtitles and narration translated into each country's language.

[1199] Privacy and Data Security

[1200] Encrypt your data to ensure the security of all your data.

[1201] Generate and preview the movie

[1202] The user can preview the generated movie and request changes to the scene order or music.

[1203] Final output and sharing

[1204] The server generates the final version of the movie and serves it to the user, who then downloads it, saves it, and shares it online.

[1205] Example prompt sentence:

[1206] "I would like to create a wedding video using the bride and groom's kiss and ring exchange scenes from the photos and videos I uploaded, and create a story with inspiring music. I would like the narration to be in both English and Spanish."

[1207] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1208] Step 1:

[1209] Data collection

[1210] Input: Photos, videos, and audio files from the user's wedding day

[1211] The user selects the digital data using a dedicated application or website and clicks an upload button.

[1212] The terminal transmits the selected data to a server via the Internet.

[1213] Output: Digital data sent to the server

[1214] Step 2:

[1215] Data analysis and classification

[1216] Input: Digital data sent to the server

[1217] The server runs a data analysis program that categorizes the photos, videos, and audio recordings into their respective categories.

[1218] The server uses image recognition technology (e.g., OpenCV, TensorFlow) to analyze the photos and video scenes and tag important moments such as the bride and groom's faces, the ring exchange, and the vow kiss.

[1219] The server uses voice recognition technology (e.g., Google Speech-to-Text API, IBM Watson) to convert the audio recording file into text and extract inspirational messages and keywords.

[1220] Output: Classified digital data, tagged key scenes, and audio data converted to text

[1221] Step 3:

[1222] Building a story

[1223] Input: Classified digital data, tagged important scenes, transcribed audio data, user preferences and themes

[1224] The server generates a storyboard using an AI model (e.g., GPT-4) based on the theme and wishes (prompt sentences) specified by the user.

[1225] The server places the tagged scenes and extracted messages on a storyboard and determines the order of each scene.

[1226] Output: Generated storyboard

[1227] Step 4:

[1228] Music and effects selection

[1229] Input: Generated storyboard

[1230] The server selects suitable songs from a music library (e.g., Epidemic Sound, Artlist) to enhance the emotion.

[1231] The server applies visual effects such as fade-in, fade-out, and slow motion to match the storyboard scenes.

[1232] Output: Story with added music and effects

[1233] Step 5:

[1234] Multilingual and culturally accommodating

[1235] Input: Stories with added music and effects, user-specified language and cultural preferences

[1236] The server uses a multilingual translation system (e.g., Google Translate API, DeepL) to translate subtitles and narration into the language specified by the user.

[1237] The server customizes the story content based on the specified culture and customs.

[1238] Output: Multilingual and culturally adapted stories

[1239] Step 6:

[1240] Privacy and Data Security

[1241] Input: Uploaded digital data

[1242] The server encrypts all digital data using the latest encryption technology (e.g., AES-256).

[1243] The server uses AI models to anonymize and mask unnecessary parts of the data to protect user privacy.

[1244] Output: Encrypted and privacy-protected digital data

[1245] Step 7:

[1246] Generate and preview the movie

[1247] Input: Multilingual and culturally adapted stories, encrypted and privacy-protected digital data

[1248] The server uses a video editing library (e.g. FFmpeg) to stitch all the elements together and generate the final movie file.

[1249] The server temporarily stores the generated movie file in cloud storage and provides a preview page URL for the user to preview it.

[1250] Users can visit a preview page to view the generated movie and provide feedback through the interface on any necessary corrections.

[1251] Output: Movie file with user review and feedback

[1252] Step 8:

[1253] Final output and sharing

[1254] Input: Movie file reviewed and feedback by user, user preferred format

[1255] The server makes corrections based on the user's feedback and generates the final version of the movie.

[1256] The server outputs the movie file in the user's desired export format (e.g., high-resolution video file, streaming link, etc.).

[1257] The user's device downloads the final movie file and stores it locally, allowing them to share it online via social media or email if desired.

[1258] Output: Final version movie file

[1259] (Application example 1)

[1260] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1261] This invention relates to a system that automatically generates moving, personalized movies simply by providing digital data (photos, videos, and audio recording files) from users. However, conventional systems lack specific application examples for further improving the user experience. In particular, they are unable to record shopping experiences in virtual stores and automatically generate movies from them, making it difficult for users to experience moving movies in different contexts. To solve this problem, it is necessary to expand the system to include movie generation based on shopping experiences in virtual stores.

[1262] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1263] In this invention, the server includes a means for uploading photos, videos, and audio recording files, a means for analyzing the uploaded data, categorizing the content, and extracting important scenes and messages, and a means for automatically generating an inspiring story based on the user's preferences and themes. This enables a shopping experience movie to be automatically generated using data taken by the user in a virtual store. The server also includes a means for performing facial recognition and tagging important moments, a means for tagging products and scenes in the virtual store, a means for converting conversation recording files into text using voice recognition technology and extracting important messages and keywords, and a means for generating inspiring prompts based on the user's shopping experience in the virtual store, further enriching the user's shopping experience and recording it as an inspiring movie.

[1264] "Photography" is a visual medium that uses light to record the shape and color of objects.

[1265] "Video" is a medium containing visual and audio information that records a time-sequence of images.

[1266] An "audio recording file" is a file that digitally records a person's voice or sound.

[1267] "Uploading" refers to the act of sending data owned by a user to a server via the Internet.

[1268] "Analysis" is the act of breaking down the provided data to understand its content and structure.

[1269] "Classification" is the act of dividing data into groups according to specific criteria.

[1270] An "important scene" is a moment that has particular meaning or impact in the overall flow or story.

[1271] A "message" is a portion of a voice recording that contains meaning or information extracted from the voice recording file.

[1272] "Automatic story generation" refers to the act of creating a coherent narrative based on analyzed data, based on the user's wishes and themes.

[1273] "Music" is sound with a melody or rhythm used to add emotional elements to a movie.

[1274] "Effects" are techniques and methods for giving a movie special visual or auditory effects.

[1275] "Multilingual support" refers to the ability to provide content in multiple languages.

[1276] "Cultural adaptation" is the act of adjusting content to suit a particular culture and customs.

[1277] "Privacy protection" refers to measures to protect personal information from unauthorized use.

[1278] "Data security" refers to the state in which data is protected from unauthorized access and tampering.

[1279] "Exporting" is the act of outputting the final movie file in a particular format.

[1280] "Sharing" is the act of making the generated movie available for viewing by others.

[1281] A "virtual store" is a virtual sales environment that offers products and services over the Internet.

[1282] "Shopping experience" is a general term for a series of actions and emotions that a user goes through when selecting and purchasing a product.

[1283] A "prompt sentence" is a sentence that is input to a generative AI model based on specified conditions and content.

[1284] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this invention, a system based on the following means is required.

[1285] 1. Data collection

[1286] The server provides a means for users to upload photos, videos, and audio recording files through dedicated applications or websites, making it easy for users to provide data.

[1287] 2. Data Analysis and Classification

[1288] The server receives the uploaded data and categorizes the photos, videos, and audio recordings into their respective categories. Image recognition technology is used to identify scenes in the photos and videos and tag specific moments. Audio recordings are converted into text using speech recognition technology, and important messages and keywords are extracted. Software such as OpenCV and TensorFlow are used for this process.

[1289] 3. Build a story

[1290] The server automatically generates the best story based on the user's specified theme and desires, key moments, and messages. It uses a generative AI model to build a storyboard and determine the order and narrative of each scene, creating a coherent and moving story.

[1291] 4. Music and Effects Selection

[1292] The server selects music from a library to enhance emotions based on the story, and adds effects appropriate for each scene. The music and effects are adjusted to match the flow of the scenes. This process is performed using software such as PIL (Python Imaging Library) and moviepy.

[1293] 5. Multilingual and culturally accommodating

[1294] The server performs multilingual text translation to generate subtitles and narration in the user's language, and customizes the content based on the user's culture and customs. By using translation software such as Google Translate API, it can accommodate users from different cultures.

[1295] 6. Privacy and Data Security

[1296] The server protects uploaded digital data using the latest encryption technology. AI models also anonymize data and mask unnecessary parts to maximize user privacy. Security software such as AES (Advanced Encryption Standard) is used.

[1297] 7. Generate and preview the movie

[1298] The server combines all the elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. Users can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[1299] 8. Final output and sharing

[1300] The server generates the final version of the movie incorporating the user's feedback and corrections. The user can choose the export format of the movie file (e.g., high-resolution video file or streaming link) according to their preference. The user receives the final movie file and can save it locally or share it on an online platform.

[1301] As a concrete example, consider a scenario where a user uploads photos and videos of clothes they tried on in a virtual store, and a shopping experience video is generated based on those photos and videos. Here is an example of a prompt to be input to the generative AI model:

[1302] "Create an inspiring shopping experience video using photos and videos of users trying on different outfits in a virtual store. At the end of the video, highlight the user's favorite outfit and play inspiring music in the background."

[1303] This prompt allows the generative AI model to generate a user experience movie that can then be previewed and shared.

[1304] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1305] Step 1:

[1306] Users upload photos, videos, and audio files through dedicated applications or websites. The input is the user's digital data, and the output is data stored on the server. Specifically, the user selects the captured data and presses the upload button, which sends the file to the server.

[1307] Step 2:

[1308] The server analyzes the uploaded data and classifies photos, videos, and audio recording files into their respective categories. The input is the uploaded digital data, and the output is the classified data. Specifically, image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are also converted into text using speech recognition technology, and important messages and keywords are extracted. This processing is done using OpenCV and TensorFlow.

[1309] Step 3:

[1310] The server references the user's specified theme and wishes, and automatically generates an optimal story based on key moments and messages. The input is classified data and the user's theme settings, and the output is a storyboard. Specifically, a generative AI model analyzes the data and constructs the story flow. This process generates a coherent and moving story.

[1311] Step 4:

[1312] Based on the story, the server selects music from a library to enhance emotions and adds appropriate effects to each scene. The input is a storyboard, and the output is the final video scene. Specifically, PIL (Python Imaging Library) and moviepy are used to insert music and effects into the video. This process creates visually and aurally appealing content.

[1313] Step 5:

[1314] The server performs multilingual text translation to generate subtitles and narration in the user's specified language. It also customizes the content based on the specified culture and customs. The input is the story text information and the user's specified language, and the output is the translated subtitles and narration. Specifically, it uses the Google Translate API or similar to translate the text data and generate the required subtitles and narration.

[1315] Step 6:

[1316] The server protects uploaded digital data using the latest encryption technology. Furthermore, the AI ​​model anonymizes the data and masks unnecessary parts to ensure maximum user privacy. The input is the user's digital data, and the output is protected data. Specifically, the data is encrypted using security technologies such as AES (Advanced Encryption Standard).

[1317] Step 7:

[1318] The server aggregates all elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. The input is the aggregated content, and the output is the generated movie file. Specifically, software such as moviepy is used to combine all scenes and effects and create the final movie file.

[1319] Step 8:

[1320] The user previews the generated movie and provides feedback through the interface on any necessary modifications (e.g., changing the order of scenes or replacing music). The input is the generated movie file and the user's feedback, and the output is the modified movie file. Specifically, the server performs the modifications based on the feedback provided by the user.

[1321] Step 9:

[1322] The server generates the final version of the movie incorporating the modifications based on the user's feedback. The export format (e.g., high-resolution video file or streaming link) can be selected according to the user's preference. The input is the modified movie file, and the output is the final movie file. Specifically, we provide the function to export the movie file in the format desired by the user.

[1323] Step 10:

[1324] Users receive the final movie file and can save it locally or share it on an online platform. The input is the final movie file, and the output is the shared movie. Specifically, the generated movie file can be uploaded to social media or cloud storage to be shared with other users.

[1325] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1326] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by the user. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more emotionally rich stories. To implement the system, it is necessary to design and implement programs based on the following means.

[1327] 1. Data collection

[1328] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. These data are temporarily stored on the device.

[1329] 2. Data transmission

[1330] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[1331] 3. Data Analysis and Classification

[1332] Server: Receives uploaded data and classifies photos, videos, and audio recordings. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recordings are also converted to text using speech recognition technology to extract important messages and keywords.

[1333] 4. Emotional Recognition

[1334] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[1335] 5. Build a story

[1336] Server: Based on the user's specified themes and wishes, and taking into account the emotional information recognized by the emotion engine, the server automatically generates a moving story. The AI ​​model determines the appropriate scene order and narrative to build a compelling story.

[1337] 6. Music and Effects Selection

[1338] Server: Based on the story, select music from the library to enhance the emotion and add effects to suit the scene, such as fade in / out and slow motion.

[1339] 7. Multilingual and culturally adaptable

[1340] Server: Generates subtitles and narration based on the user-specified language. Uses text translation tools for multilingual support and adjusts content for cultural adaptation.

[1341] 8. Privacy and Data Security

[1342] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1343] 9. Generate and preview the movie

[1344] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and users can view it using the preview link.

[1345] Users: can use the preview link to see the generated movie and provide feedback, such as requests for changes to the scene order or music.

[1346] 10. Receiving and Responding to Feedback

[1347] Server: Receives user feedback and modifies the movie based on the instructions. The modified movie is then regenerated and saved.

[1348] 11. Final output and sharing

[1349] Server: Generates the final version of the movie and provides it to the user in the export format of their choice, such as a high-resolution video file or a streaming link.

[1350] Device: Receive the final movie file and save it locally or share it on an online platform.

[1351] Specific examples

[1352] For example, consider the case of creating a wedding movie.

[1353] 1. Data collection

[1354] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[1355] 2. Data transmission

[1356] Terminal: Sends these uploaded data to the server.

[1357] 3. Data Analysis and Classification

[1358] Server: Recognizes and tags important moments in photos and videos (e.g., the promise kiss, the ring exchange), and extracts inspirational messages from audio recordings.

[1359] 4. Emotional Recognition

[1360] Server: Analyzes facial expressions and tone of voice from video and audio to recognize the emotions of the bride and groom and guests. For example, it identifies touching moments by detecting smiles and tears.

[1361] 5. Build a story

[1362] Server: Based on the extracted scenes and messages and the emotional information recognized by the emotion engine, a moving story of the entire wedding is constructed.

[1363] 6. Music and Effects Selection

[1364] Server: Add inspiring music and apply effects appropriate to the scene, including fade in / out, slow motion, etc.

[1365] 7. Multilingual and culturally adaptable

[1366] Server: For international weddings, generate subtitles and narration translated into each country's language.

[1367] 8. Privacy and Data Security

[1368] Server: All data is encrypted and secured.

[1369] 9. Generate and preview the movie

[1370] User: Preview the generated movie and request corrections if necessary.

[1371] 10. Receiving and Responding to Feedback

[1372] Server: Based on user feedback, the movie is regenerated and modifications are made.

[1373] 11. Final output and sharing

[1374] Server: Generates the final version of the movie and serves it to the user.

[1375] Device: Download movies and finally save and share them.

[1376] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[1377] The processing flow will be explained below.

[1378] Step 1:

[1379] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. Uploaded data is temporarily stored on the device.

[1380] Step 2:

[1381] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[1382] Step 3:

[1383] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[1384] Step 4:

[1385] Server: Analyzes photo data using image recognition technology, automatically identifying faces and tagging distinctive scenes (e.g., smiles or touching moments).

[1386] Step 5:

[1387] Server: Analyzes video data using the video analysis module, detects scene changes, and extracts important moments (e.g., the engagement kiss or ring exchange).

[1388] Step 6:

[1389] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[1390] Step 7:

[1391] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[1392] Step 8:

[1393] Server: Automatically generates an emotional storyboard based on the user's specified themes and wishes, and also takes into account recognized emotional information. The AI ​​model determines the order of scenes and the progression of the story, building a coherent story.

[1394] Step 9:

[1395] Server: Based on the story, select music from the library to enhance the emotions and add effects to each scene, such as fade in / out and slow motion.

[1396] Step 10:

[1397] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content to take cultural adaptation into account.

[1398] Step 11:

[1399] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1400] Step 12:

[1401] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and a preview link is provided to users.

[1402] Step 13:

[1403] Users: Use the preview link to view the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[1404] Step 14:

[1405] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[1406] Step 15:

[1407] Server: Generates the final movie file and exports it in the format of your choice (e.g. high-resolution video file or streaming link).

[1408] Step 16:

[1409] Terminal: Receive the final movie file and save it locally. Furthermore, if you want to share it on an online platform, you can easily do so using a dedicated application.

[1410] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[1411] Example 2

[1412] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1413] In recent years, the widespread use of smartphones and digital cameras has increased the opportunities for users to save various events and everyday moments as digital data (photos, videos, and audio recording files). However, editing this data to create moving and personalized movies requires advanced editing skills and time, making it difficult for average users. Conventional methods have been problematic in that the editing process is cumbersome and it is difficult to automatically generate emotionally appealing stories, leaving users unable to obtain satisfactory results.

[1414] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to upload photos, videos, and audio recording files; a means for transmitting the uploaded data to the server via the Internet; a means for the server to analyze the received data, classify the content, and extract important scenes and messages; a means for analyzing the user's emotions using emotion recognition technology; a means for automatically generating a story based on the user's preferences and themes, taking into account the emotional information; a means for adding music and effects to the story; a means for multilingual support and cultural adaptation; a means for presenting the generated movie to the user and making corrections based on the user's feedback; a means for ensuring privacy protection and data security; and a means for exporting and sharing the final movie file in a specified format. This allows even ordinary users to easily automatically generate moving and personalized movies and obtain high-quality results.

[1415] "User" refers to any individual or organization that uses this system.

[1416] "Photo" refers to still image data that a user takes using a digital device and uploads to the system.

[1417] "Video" refers to moving image data that a user shoots using a digital device and uploads to the system.

[1418] "Audio recording file" refers to a file that contains audio data recorded by a user or another person.

[1419] "Upload" refers to a user sending a photo, video, or audio recording file to the system through a dedicated application or website.

[1420] "Internet" refers to a global network for data communication.

[1421] "Server" refers to a computer system that receives, analyzes, sorts, and processes data sent by users.

[1422] "Analysis" refers to the process by which the server understands and interprets the content of photographs, videos, and audio recordings based on data.

[1423] "Classification" refers to the process by which the server divides the data it receives into specific categories.

[1424] "Important scenes and messages" refer to notable moments or words that the user finds interesting or that the system automatically selects.

[1425] "Emotion recognition technology" refers to the technology that a system uses to analyze and identify a user's emotions from images and audio.

[1426] "Automatic story generation" refers to the process in which a system constructs a series of events into a narrative based on the user's wishes, themes, and emotional information.

[1427] "Music and Effects" refers to background music and visual effects added to enhance the visual and auditory impression of the movie.

[1428] "Multilingual" refers to the ability to display content or provide narration in different languages.

[1429] "Cultural adaptation" refers to adjusting content to suit the user's culture.

[1430] "Feedback" refers to the act of a user conveying to the system their opinions and requests for corrections and improvements to a movie.

[1431] "Privacy protection" refers to protecting users' personal information and data from being leaked to third parties.

[1432] "Data security" refers to measures taken to keep data protected from unauthorized access and tampering.

[1433] "Export" refers to the act of saving or outputting the final movie file in a particular format.

[1434] "Sharing" refers to sharing the generated movie file or link with other users and platforms.

[1435] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. Specific procedures for implementing the present invention and the hardware and software used are described below.

[1436] This system is primarily composed of users, devices, and a server. Users upload data using a dedicated application or website, and their devices send the data to the server via the Internet. The server analyzes the data, generates a story, and provides the final movie file to the user. The specific operation is explained below.

[1437] Data collection

[1438] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website, and these data are temporarily stored on the device.

[1439] Sending data

[1440] The device sends the data uploaded by the user to a server via the Internet, including metadata (e.g., date, time, location, etc.).

[1441] Data analysis and classification

[1442] The server analyzes the received data and classifies each photo, video, and audio recording file using image recognition technologies (e.g., Google Cloud Vision API) and speech recognition technologies (e.g., Google Cloud Speech-to-Text). Scenes in the photos and videos are identified and specific moments (e.g., the engagement kiss or the ring exchange) are tagged. Audio recordings are converted to text and key messages and keywords are extracted.

[1443] Emotion recognition

[1444] The server uses emotion recognition technology (e.g., Microsoft Azure Emotion API) to recognize the user's emotions from video and audio, analyzing smiles, tears, tone of voice, etc. to identify emotional moments.

[1445] Building a story

[1446] The server automatically generates a moving story based on the user's preferences and themes, taking into account emotional information. This story generation uses an AI model (e.g., OpenAI GPT) to determine the order of movie scenes and the content of the narration.

[1447] Music and effects selection

[1448] Based on the story, the server selects music from a library (e.g., Epidemic Sound) to enhance the emotions and adds effects (fade in / out, slow motion, etc.) that match the scene.

[1449] Multilingual and culturally accommodating

[1450] The server generates subtitles and narration based on the language specified by the user, and uses text translation tools (e.g., Google Translate API) to support multiple languages, adjusting the content for cultural adaptation.

[1451] Privacy and Data Security

[1452] The server uses the latest encryption technology (e.g., AES-256) for all uploaded data to ensure privacy and data security. Data is sent and received using SSL / TLS encryption, and data is encrypted when stored.

[1453] Generate and preview the movie

[1454] The server combines all the elements to generate a first version of the movie file, stores it temporarily, and makes it available to users via a preview link.

[1455] Users can use the preview link to view the resulting movie and submit requests for changes to the scene order or music, if necessary.

[1456] Receiving and responding to feedback

[1457] The server receives feedback from the user and modifies the movie based on the instructions, and the modified movie is regenerated and saved.

[1458] Final output and sharing

[1459] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link.

[1460] The device will receive the final movie file and store it locally, or you can share it on an online platform.

[1461] Examples of prompt statements

[1462] Below are some example prompts to input to the generative AI model:

[1463] Create a moving wedding video using photos and videos from your wedding, along with messages from friends and family. Highlight the bride and groom's smiling and tearful moments, and add music appropriate to their special day. For an international wedding, add subtitles in English and Japanese to reflect cultural elements. Specifically, use emotional music for the groom's entrance and select a song that gradually builds up during the bride's entrance. Additionally, highlight the vows and ring exchange in slow motion, and add moving guest messages.

[1464] Through the above steps, the present invention utilizes an emotion engine to automatically generate an even more moving movie based on digital data that records a user's special moments, and provides it as a personalized keepsake.

[1465] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1466] Step 1: Collect data

[1467] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The input is the user's own digital data. The uploaded data is temporarily stored on the device and becomes the output. For example, a user can drag and drop photos from their wedding day or audio files containing messages from friends into the application, and then enter the event theme and desired details into a form.

[1468] Step 2: Sending data

[1469] The device sends the data uploaded by the user to the server via the Internet. The input includes data temporarily stored on the device and metadata (date, time, location, etc.). The output is data sent to the server and received by the server. Specifically, the device securely sends this data to the server using the HTTPS protocol.

[1470] Step 3: Analyze and classify the data

[1471] The server analyzes and classifies the data it receives. The input is the data received by the server. The output is photos, videos, and audio recordings sorted into their respective categories, with each file tagged with important scenes and keywords. For example, the server uses Google Cloud Vision API to identify scenes in photos and videos and tag specific moments, such as the bride and groom's kiss and ring exchange. It also uses Google Cloud Speech-to-Text to convert audio recordings into text and extract important messages and keywords.

[1472] Step 4: Recognize emotions

[1473] The server uses emotion recognition technology to analyze user emotions from video and audio. The input is analyzed and classified video and audio data. The output is an emotional tag attached to each piece of data. Specifically, it uses the Microsoft Azure Emotion API to analyze smiles, tears, tone of voice, and other emotions to identify moments when the bride and groom and guests are emotional.

[1474] Step 5: Build your story

[1475] The server automatically generates a story based on the user's wishes and themes, taking into account emotional information. The inputs are data tagged with emotions and the wishes and themes provided by the user. The output is a movie plan that determines the order of scenes and narration for a moving story. Specifically, it uses OpenAI GPT to construct the story's narrative and determine the placement of scenes.

[1476] Step 6: Choose music and effects

[1477] The server selects music and effects based on the story. The input is a story plan. The output is a story plan with music and effects applied to scenes to enhance the emotions. Specific operations include selecting appropriate music from the Epidemic Sound library and applying effects such as fade-in / out and slow motion to scenes.

[1478] Step 7: Multilingualism and cultural adaptation

[1479] The server generates subtitles and narration based on the language specified by the user. The inputs are the language information specified by the user and a story plan. The output is a story plan that includes subtitles and narration in multiple languages. Specifically, the system uses the Google Translate API to translate the text into multiple languages ​​and edits the content to take cultural adaptation into account.

[1480] Step 8: Privacy and Data Security

[1481] The server uses encryption technology when sending, receiving, and storing data. The inputs include user data stored on the server and data being sent and received. The output is encrypted data. Specifically, data security is ensured using AES-256 data encryption and SSL / TLS protocols.

[1482] Step 9: Generate and preview your movie

[1483] The server integrates all elements to generate a first-run movie file and temporarily stores it. The input is the final story plan. The output is the first-run movie file. Specifically, video editing software is used to generate the movie based on the pre-planned layout.

[1484] The user uses the preview link to view the generated movie. The input is the generated movie file. The output is feedback based on the preview results.

[1485] Step 10: Receiving and responding to feedback

[1486] The server receives user feedback and modifies the movie based on the instructions. The inputs are the user's feedback and the original movie file. The output is a modified movie file that reflects the feedback. Specifically, the server reorders the requested scenes and adjusts the music to generate a new movie.

[1487] Step 11: Final output and sharing

[1488] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link. The input is the final, modified movie plan. The output is the final movie file. Specific operations include generating an HD MP4 file or a YouTube streaming link as the final output.

[1489] The device receives the final movie file and stores it locally. The input is the final movie file provided by the server. The output is a locally stored movie file, which the user can also upload to an online platform for sharing.

[1490] (Application example 2)

[1491] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1492] Conventional digital content generation systems that use photos, videos, and audio files have struggled to automatically generate emotionally rich stories. Furthermore, there are challenges, such as generating movies that effectively reflect the user's emotions, supporting multiple languages, and protecting privacy. In particular, there is a need for a system that can visually and emotionally recreate users' memories and share them in high resolution with simple smartphone application operations.

[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, and means for automatically generating an emotional story based on the user's wishes and themes. This makes it possible to automatically generate an emotional digital story that effectively reflects the user's emotions using an emotion engine.

[1494] A "photograph" is a file that digitally records a still image.

[1495] A "video" is a file that contains a digital recording of moving images.

[1496] An "audio recording file" is a file that records audio in digital format.

[1497] "Uploading" means sending local data owned by the user to a server via the Internet.

[1498] "Data analysis" refers to the use of machine learning and algorithms to process uploaded digital content and extract specific information.

[1499] "Content classification" refers to grouping photos, videos, and audio recording files based on their format and characteristics.

[1500] "Key Scene and Message Extraction" refers to automatically identifying moments and meaningful messages within digital content that deserve special emphasis.

[1501] "User's wishes and themes" are requests and objectives for a story or movie specified by the user.

[1502] "Automatic generation of emotional stories" is the process of automatically constructing emotional stories based on extracted scenes and messages.

[1503] An "emotion engine" is an algorithm that recognizes user emotions from digital data and uses that information to create stories and movies.

[1504] "Adding Music and Effects" is the process of adding music and visual effects to a moving story to enhance its emotional expression.

[1505] "Multilingual support" means translating generated content into multiple languages ​​to accommodate users from different cultures and regions.

[1506] "Cultural adaptation" means adjusting content to suit users in a particular culture or region.

[1507] "Privacy protection" refers to measures and technologies to protect user data and information from third parties.

[1508] "Data security" means ensuring that digital data is protected and that unauthorized access is prevented.

[1509] "Final Movie File" is the finished digital movie file exported in the user's desired format.

[1510] "Export in specified format" means outputting the final movie in a file format selected by the user.

[1511] "Emotion analysis" refers to recognizing and identifying a user's emotional state from digital data.

[1512] "Previewing a movie" means checking a movie before it is completed.

[1513] "Accepting a change request" means receiving a request for correction or addition from a user and changing the content.

[1514] A "high resolution movie file" is a digital movie file with high quality video output.

[1515] "Share on social media" means posting and sharing the generated movie file on an internet communication platform.

[1516] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. By combining this system with an emotion engine, it creates a story that reflects the user's emotions.

[1517] Program Generation and Processing Description

[1518] Hardware and software used:

[1519] Server: High performance data processing and storage server.

[1520] Terminal (smartphone or tablet): A device through which a user uploads digital data.

[1521] Network Connection: An internet connection for uploading, downloading, and communicating data.

[1522] Software used:

[1523] TensorFlow: A machine learning framework for image and speech recognition.

[1524] OpenCV: A library for image and video analysis.

[1525] Django: A server-side web framework.

[1526] React Native: A development platform for smartphone applications.

[1527] Amazon Web Services (AWS) S3: Data storage service.

[1528] Amazon Polly: A speech synthesis service.

[1529] Examples of implementation:

[1530] 1. Uploading data:

[1531] Users use a smartphone app to upload photos, videos, and audio recordings related to events such as weddings and birthdays, and metadata (date, time, and location information) is automatically captured and added to the uploaded data.

[1532] 2. Data transmission:

[1533] The device sends the data uploaded by the user to a server via the Internet, where it is stored in Amazon Web Services (AWS) S3.

[1534] 3. Data analysis and classification:

[1535] The server receives uploaded data using the Django framework and leverages TensorFlow and OpenCV to analyze and classify photo, video, and audio recording files, automatically extracting and tagging important scenes and messages from the uploaded data.

[1536] 4. Emotion Recognition:

[1537] The server uses an emotion engine to analyze emotions in the data, for example by analyzing facial expressions and tone of voice in the video to identify moments of emotion or joy.

[1538] 5. Build a story:

[1539] The server automatically generates an emotional story based on the user's specified theme and wishes, and this process incorporates the emotional information recognized by the emotion engine.

[1540] 6. Add music and effects:

[1541] The server adds appropriate music and effects to the generated story, and also generates voice narration using Amazon Polly, including multilingual support and cultural adaptation.

[1542] 7. Generate and preview the movie:

[1543] The server generates the initial movie file, allows users to preview it through a smartphone app, and accepts and responds to user requests for changes.

[1544] 8. Final movie output:

[1545] Users can download the final version of their movie in high resolution and share it on social media, all while ensuring privacy and data security.

[1546] Examples of prompts:

[1547] "Analyze family trip data to create an inspiring movie that captures moments of happiness for users."

[1548] In this way, the system of the present invention utilizes digital data of special moments provided by the user to automatically generate moving and personalized movies, allowing memories to be recreated in a deeper and more emotional way.

[1549] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1550] Step 1:

[1551] Data collection and upload

[1552] Users use a smartphone application to upload photos, videos, and audio recordings. Specifically, they tap the app's "upload" button and select media files from their device. The input data are photos, videos, and audio recordings and their associated metadata. The output data is digital data sent to a server.

[1553] Step 2:

[1554] Sending data

[1555] The device sends the data uploaded by the user to a server via the Internet. Specifically, the application transfers the selected file to the server using the HTTPS protocol. The input data is the digital content specified by the user through the application. The output data is the file stored on the server.

[1556] Step 3:

[1557] Data analysis and classification

[1558] The server receives, analyzes, and classifies the uploaded data. It uses TensorFlow and OpenCV to analyze photo, video, and audio recording files and processes the data according to their respective formats. Specific operations include scene analysis using image recognition and text conversion using voice recognition. The input data is the uploaded digital data. The output data is the classified content and its metadata.

[1559] Step 4:

[1560] Emotion recognition

[1561] The server uses an emotion engine to analyze users' emotions from uploaded videos and audio. Specifically, it analyzes facial expressions and tone of voice to identify emotional elements and emotional moments. The input data are classified photos, videos, and audio recordings. The output data is media data with recognized emotion information.

[1562] Step 5:

[1563] Building a story

[1564] The server automatically generates an inspiring story based on the user's specified theme and wishes. It primarily uses an AI model to construct the story based on recognized emotional information and tagged scenes. Specifically, it determines the order of scenes and automatically forms a narrative. The input data is media data with emotional information and the user's specified theme. The output data is the constructed story.

[1565] Step 6:

[1566] Adding music and effects

[1567] The server adds appropriate music and visual effects to the generated story. It uses Amazon Polly to generate narration, and also supports multiple languages ​​and cultural adaptations. Specific operations include selecting music, applying effects, and inserting narration. The input data is the constructed story. The output data is a movie with music and effects added.

[1568] Step 7:

[1569] Generate and preview the movie

[1570] The server generates the final movie file and allows users to preview it through a smartphone app. Specific operations include rendering the movie file and generating a preview file. The input data is movie data with music and effects added. The output data is a movie file that users can preview.

[1571] Step 8:

[1572] Receiving and responding to feedback

[1573] The user previews the movie generated by the smartphone app and sends correction requests as necessary. Specifically, the user uses the feedback form on the preview screen to enter changes. The input data is the user's feedback. The output data is the feedback sent to the server as a correction request.

[1574] Step 9:

[1575] Final output and sharing

[1576] The server generates a final movie file incorporating the user's feedback and provides a download link. Specific operations include rendering the final movie file, generating a high-resolution file, and emailing the download link. The input data is the movie data with correction instructions. The output data is the final high-resolution movie file and a download link.

[1577] Through the above processing steps, users can easily create, save, and share moving and personalized movies based on digital data recordings of special moments.

[1578] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1579] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1580] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1581] [Fourth embodiment]

[1582] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1583] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1584] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1585] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1586] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1587] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1588] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1589] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1590] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1591] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1592] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1593] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1594] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1595] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, it is necessary to design and implement programs based on the following means.

[1596] 1. Data collection

[1597] Users upload photos, videos, and audio files to the system through dedicated applications or websites. The digital data provided is then sent to the server.

[1598] 2. Data Analysis and Classification

[1599] Server: Receives uploaded data and categorizes photos, videos, and audio recording files into their respective categories. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are converted to text using speech recognition technology to extract important messages and keywords.

[1600] 3. Build a story

[1601] Server: Based on the user's specified theme and wishes, the server automatically generates the optimal story based on key moments and messages. The AI ​​model builds the storyboard and determines the order and narrative of each scene.

[1602] 4. Music and Effects Selection

[1603] Server: Based on the story, select music from the library to enhance emotions and add effects appropriate for each scene. The music and effects are adjusted to fit the flow of the scene.

[1604] 5. Multilingual and culturally accommodating

[1605] Server: Performs multilingual text translation to generate subtitles and narration in the user's language, and customizes content based on the user's culture and customs.

[1606] 6. Privacy and Data Security

[1607] Server: Uploaded digital data is protected with the latest encryption technology, and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[1608] 7. Generate and preview the movie

[1609] Server: Integrates all elements and generates the final movie file, which is temporarily stored and made available for user preview.

[1610] User: Can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[1611] 8. Final output and sharing

[1612] Server: Generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g. high-resolution video file or streaming link) can be selected based on the user's preference.

[1613] Device: Receive the final movie file and save it locally or share it on an online platform.

[1614] Specific examples

[1615] For example, consider the case of creating a wedding movie.

[1616] 1. Data collection

[1617] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[1618] 2. Data Analysis and Classification

[1619] Server: Analyzes the bride and groom's appearance in photos and tags the moments of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes recorded conversations to extract moving messages.

[1620] 3. Build a story

[1621] Server: Arrange individual photos and video scenes on a storyboard to build a moving story from the bride and groom meeting to their wedding.

[1622] 4. Music and Effects Selection

[1623] Server: Insert moving music and apply effects (fade in / out, slow motion, etc.) to suit the scene.

[1624] 5. Multilingual and culturally accommodating

[1625] Server: For international weddings, generate subtitles and narration translated into each country's language.

[1626] 6. Privacy and Data Security

[1627] Server: All data is encrypted and secured.

[1628] 7. Generate and preview the movie

[1629] User: Preview the generated movie and request changes to the scene order or music.

[1630] 8. Final output and sharing

[1631] Server: Generates the final version of the movie and serves it to the user.

[1632] Device: Download movies, save them, and share them online.

[1633] In this way, the system according to the present invention automatically generates an inspiring movie based on the digital data recording a special moment of the user, and provides it as a personalized keepsake.

[1634] The processing flow will be explained below.

[1635] Step 1:

[1636] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The data is temporarily stored on the device.

[1637] Step 2:

[1638] Terminal: The data uploaded by the user is sent to the server via the Internet, including metadata (such as date, time, and location information).

[1639] Step 3:

[1640] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[1641] Step 4:

[1642] Server: Analyzes photo data using image recognition technology, automatically tags faces and distinctive scenes (e.g., smiles, touching moments).

[1643] Step 5:

[1644] Server: The video analysis module analyzes the video data, detects scene changes, and extracts important moments, such as the engagement kiss and ring exchange.

[1645] Step 6:

[1646] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[1647] Step 7:

[1648] Server: Generates a storyboard using key scenes and messages based on the user's specified theme and wishes. The AI ​​model determines the order of scenes and narrative progression to build a compelling story.

[1649] Step 8:

[1650] Server: Select music from the library to enhance the emotions of the story and add effects to each scene, including fade-in / out and slow motion.

[1651] Step 9:

[1652] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content for cultural adaptation.

[1653] Step 10:

[1654] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1655] Step 11:

[1656] Server: Integrates all elements and generates the initial movie version. The generated movie is temporarily saved and a preview link is provided to the user.

[1657] Step 12:

[1658] Users: Use the preview link to see the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[1659] Step 13:

[1660] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[1661] Step 14:

[1662] Server: Generates the final version of your movie and exports it in the format of your choice, including high-resolution video files and streaming links.

[1663] Step 15:

[1664] Terminal: Receive the final movie file and save it locally. If you want to share it on an online platform, you can easily do so using a dedicated application.

[1665] These are the basic processing steps in the embodiment of the present invention, which allows users to effectively record special moments and easily share them as moving movies.

[1666] Example 1

[1667] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1668] There is a need for a system that allows users to easily create inspiring and personalized movies. It is also necessary to meet the needs of each user for privacy protection, data security, multilingual support, and cultural adaptation. Conventional systems cannot comprehensively provide these features, and they are time-consuming and labor-intensive.

[1669] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1670] In this invention, the server includes means for users to upload photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, means for automatically generating an inspiring story based on the user's wishes and themes, means for adding music and effects necessary for the inspiring story, means for providing multilingual support and cultural adaptation to the generated story, means for presenting the generated movie to the user and making corrections based on feedback, means for encrypting digital data to ensure privacy protection and data security, and means for exporting and sharing the final movie file in a specified format, thereby enabling users to easily create inspiring and personalized movies that also provide privacy protection, multilingual support, and cultural adaptation.

[1671] "Uploading" refers to the transfer of digital data from a user's terminal to a remote system such as a server.

[1672] "Data analysis" is the process of analyzing uploaded digital data, categorizing it into different formats such as photos, videos, and audio recording files, and understanding its content.

[1673] "Content classification" refers to the act of organizing analyzed digital data based on its type and characteristics, separating it into categories such as photos, videos, and audio recording files.

[1674] "Significant scene and message extraction" is the process of identifying and tagging particularly meaningful scenes and messages within photographs, videos, and audio recordings.

[1675] "Means for automated story generation" refers to algorithms and AI models that create compelling narrative structures from digital data based on user preferences and themes.

[1676] "Adding music and effects" is the process of applying emotionally enhancing music and visual effects to the generated story.

[1677] "Multilingual support" refers to the ability to translate generated content into multiple user-specified languages ​​and include subtitles and narration.

[1678] "Cultural adaptation" is the process of appropriately customizing content based on the user's specified culture and customs.

[1679] "Privacy protection" refers to various measures taken to protect users' personal information and digital data from others.

[1680] "Data security" means the technological measures and processes used to ensure the confidentiality, integrity and protection from unauthorized access of uploaded digital data.

[1681] "Export" refers to the act of converting the final generated movie file into a specific format (e.g., MP4, AVI, etc.) and outputting it.

[1682] "Sharing" refers to providing the generated movie file to other users or platforms so that they can access it.

[1683] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this system, the following hardware and software are used:

[1684] 1. Hardware and Software

[1685] Hardware: Servers, user devices (PCs, smartphones, etc.)

[1686] Software: Uploading applications and websites, image recognition technology (e.g., OpenCV, TensorFlow), speech recognition technology (e.g., Google Speech-to-Text API, IBM Watson), multilingual text translation systems (e.g., Google Translate API, DeepL), AI models (e.g., GPT-4)

[1687] 2. Detailed processing instructions

[1688] 1. Data collection

[1689] Users upload digital data (e.g., photos, videos, and audio files from the wedding day) through a dedicated application or website, and this data is sent to the server.

[1690] 2. Data Analysis and Classification

[1691] The server receives the uploaded data and categorizes it into photos, videos, and audio recordings. Image recognition technology is used to identify scenes in the photos and videos, confirm the bride and groom's appearance, and tag moments such as the vow kiss and ring exchange. Audio files are converted to text using speech recognition technology, which extracts important messages and keywords.

[1692] 3. Build a story

[1693] The server automatically generates the optimal story based on the user's specified theme and wishes, key moments, and messages. It uses a generative AI model (e.g., GPT-4) to build the storyboard and determine the order and narrative of each scene.

[1694] 4. Music and Effects Selection

[1695] The server selects inspiring music from a library and adds appropriate effects for each scene, with the music and effects tailored to the flow of the scene.

[1696] 5. Multilingual and culturally accommodating

[1697] The server uses a multilingual translation system to generate subtitles and narration in the language specified by the user, and also customizes the content based on the specified culture and customs.

[1698] 6. Privacy and Data Security

[1699] The server protects uploaded digital data using the latest encryption technology (e.g., AES-256), and AI models anonymize and mask unnecessary parts of the data to maximize user privacy.

[1700] 7. Generate and preview the movie

[1701] The server aggregates all the elements and generates the final movie file using a video editing library (e.g., FFmpeg), which is temporarily stored in cloud storage and made available for users to preview.

[1702] 8. Final output and sharing

[1703] The server generates the final version of the movie incorporating the user's feedback and corrections. The export format (e.g., high-resolution video file or streaming link) can be selected based on the user's preference. The user's device receives the final movie file and can save it locally or share it on an online platform.

[1704] 3. Examples of concrete examples and prompts

[1705] A specific example of creating a wedding movie will be given below.

[1706] Data collection

[1707] Users upload photos and videos from their wedding day, as well as recorded messages from friends and family.

[1708] Data analysis and classification

[1709] The server analyzes the bride and groom's appearance in photos and tags the moment of the vow kiss and ring exchange, extracts important scenes from videos, and transcribes audio messages to extract moving messages.

[1710] Building a story

[1711] The server uses these photos and video footage to build a moving story from the moment the bride and groom met to their wedding day.

[1712] Music and effects selection

[1713] Choose inspiring music and add effects that suit the scene (e.g. fade in / out, slow motion, etc.).

[1714] Multilingual and culturally accommodating

[1715] For international weddings, generate subtitles and narration translated into each country's language.

[1716] Privacy and Data Security

[1717] Encrypt your data to ensure the security of all your data.

[1718] Generate and preview the movie

[1719] The user can preview the generated movie and request changes to the scene order or music.

[1720] Final output and sharing

[1721] The server generates the final version of the movie and serves it to the user, who then downloads it, saves it, and shares it online.

[1722] Example prompt sentence:

[1723] "I would like to create a wedding video using the bride and groom's kiss and ring exchange scenes from the photos and videos I uploaded, and create a story with inspiring music. I would like the narration to be in both English and Spanish."

[1724] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1725] Step 1:

[1726] Data collection

[1727] Input: Photos, videos, and audio files from the user's wedding day

[1728] The user selects the digital data using a dedicated application or website and clicks an upload button.

[1729] The terminal transmits the selected data to a server via the Internet.

[1730] Output: Digital data sent to the server

[1731] Step 2:

[1732] Data analysis and classification

[1733] Input: Digital data sent to the server

[1734] The server runs a data analysis program that categorizes the photos, videos, and audio recordings into their respective categories.

[1735] The server uses image recognition technology (e.g., OpenCV, TensorFlow) to analyze the photos and video scenes and tag important moments such as the bride and groom's faces, the ring exchange, and the vow kiss.

[1736] The server uses voice recognition technology (e.g., Google Speech-to-Text API, IBM Watson) to convert the audio recording file into text and extract inspirational messages and keywords.

[1737] Output: Classified digital data, tagged key scenes, and audio data converted to text

[1738] Step 3:

[1739] Building a story

[1740] Input: Classified digital data, tagged important scenes, transcribed audio data, user preferences and themes

[1741] The server generates a storyboard using an AI model (e.g., GPT-4) based on the theme and wishes (prompt sentences) specified by the user.

[1742] The server places the tagged scenes and extracted messages on a storyboard and determines the order of each scene.

[1743] Output: Generated storyboard

[1744] Step 4:

[1745] Music and effects selection

[1746] Input: Generated storyboard

[1747] The server selects suitable songs from a music library (e.g., Epidemic Sound, Artlist) to enhance the emotion.

[1748] The server applies visual effects such as fade-in, fade-out, and slow motion to match the storyboard scenes.

[1749] Output: Story with added music and effects

[1750] Step 5:

[1751] Multilingual and culturally accommodating

[1752] Input: Stories with added music and effects, user-specified language and cultural preferences

[1753] The server uses a multilingual translation system (e.g., Google Translate API, DeepL) to translate subtitles and narration into the language specified by the user.

[1754] The server customizes the story content based on the specified culture and customs.

[1755] Output: Multilingual and culturally adapted stories

[1756] Step 6:

[1757] Privacy and Data Security

[1758] Input: Uploaded digital data

[1759] The server encrypts all digital data using the latest encryption technology (e.g., AES-256).

[1760] The server uses AI models to anonymize and mask unnecessary parts of the data to protect user privacy.

[1761] Output: Encrypted and privacy-protected digital data

[1762] Step 7:

[1763] Generate and preview the movie

[1764] Input: Multilingual and culturally adapted stories, encrypted and privacy-protected digital data

[1765] The server uses a video editing library (e.g. FFmpeg) to stitch all the elements together and generate the final movie file.

[1766] The server temporarily stores the generated movie file in cloud storage and provides a preview page URL for the user to preview it.

[1767] Users can visit a preview page to view the generated movie and provide feedback through the interface on any necessary corrections.

[1768] Output: Movie file with user review and feedback

[1769] Step 8:

[1770] Final output and sharing

[1771] Input: Movie file reviewed and feedback by user, user preferred format

[1772] The server makes corrections based on the user's feedback and generates the final version of the movie.

[1773] The server outputs the movie file in the user's desired export format (e.g., high-resolution video file, streaming link, etc.).

[1774] The user's device downloads the final movie file and stores it locally, allowing them to share it online via social media or email if desired.

[1775] Output: Final version movie file

[1776] (Application example 1)

[1777] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1778] This invention relates to a system that automatically generates moving, personalized movies simply by providing digital data (photos, videos, and audio recording files) from users. However, conventional systems lack specific application examples for further improving the user experience. In particular, they are unable to record shopping experiences in virtual stores and automatically generate movies from them, making it difficult for users to experience moving movies in different contexts. To solve this problem, it is necessary to expand the system to include movie generation based on shopping experiences in virtual stores.

[1779] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1780] In this invention, the server includes a means for uploading photos, videos, and audio recording files, a means for analyzing the uploaded data, categorizing the content, and extracting important scenes and messages, and a means for automatically generating an inspiring story based on the user's preferences and themes. This enables a shopping experience movie to be automatically generated using data taken by the user in a virtual store. The server also includes a means for performing facial recognition and tagging important moments, a means for tagging products and scenes in the virtual store, a means for converting conversation recording files into text using voice recognition technology and extracting important messages and keywords, and a means for generating inspiring prompts based on the user's shopping experience in the virtual store, further enriching the user's shopping experience and recording it as an inspiring movie.

[1781] "Photography" is a visual medium that uses light to record the shape and color of objects.

[1782] "Video" is a medium containing visual and audio information that records a time-sequence of images.

[1783] An "audio recording file" is a file that digitally records a person's voice or sound.

[1784] "Uploading" refers to the act of sending data owned by a user to a server via the Internet.

[1785] "Analysis" is the act of breaking down the provided data to understand its content and structure.

[1786] "Classification" is the act of dividing data into groups according to specific criteria.

[1787] An "important scene" is a moment that has particular meaning or impact in the overall flow or story.

[1788] A "message" is a portion of a voice recording that contains meaning or information extracted from the voice recording file.

[1789] "Automatic story generation" refers to the act of creating a coherent narrative based on analyzed data, based on the user's wishes and themes.

[1790] "Music" is sound with a melody or rhythm used to add emotional elements to a movie.

[1791] "Effects" are techniques and methods for giving a movie special visual or auditory effects.

[1792] "Multilingual support" refers to the ability to provide content in multiple languages.

[1793] "Cultural adaptation" is the act of adjusting content to suit a particular culture and customs.

[1794] "Privacy protection" refers to measures to protect personal information from unauthorized use.

[1795] "Data security" refers to the state in which data is protected from unauthorized access and tampering.

[1796] "Exporting" is the act of outputting the final movie file in a particular format.

[1797] "Sharing" is the act of making the generated movie available for viewing by others.

[1798] A "virtual store" is a virtual sales environment that offers products and services over the Internet.

[1799] "Shopping experience" is a general term for a series of actions and emotions that a user goes through when selecting and purchasing a product.

[1800] A "prompt sentence" is a sentence that is input to a generative AI model based on specified conditions and content.

[1801] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. To implement this invention, a system based on the following means is required.

[1802] 1. Data collection

[1803] The server provides a means for users to upload photos, videos, and audio recording files through dedicated applications or websites, making it easy for users to provide data.

[1804] 2. Data Analysis and Classification

[1805] The server receives the uploaded data and categorizes the photos, videos, and audio recordings into their respective categories. Image recognition technology is used to identify scenes in the photos and videos and tag specific moments. Audio recordings are converted into text using speech recognition technology, and important messages and keywords are extracted. Software such as OpenCV and TensorFlow are used for this process.

[1806] 3. Build a story

[1807] The server automatically generates the best story based on the user's specified theme and desires, key moments, and messages. It uses a generative AI model to build a storyboard and determine the order and narrative of each scene, creating a coherent and moving story.

[1808] 4. Music and Effects Selection

[1809] The server selects music from a library to enhance emotions based on the story, and adds effects appropriate for each scene. The music and effects are adjusted to match the flow of the scenes. This process is performed using software such as PIL (Python Imaging Library) and moviepy.

[1810] 5. Multilingual and culturally accommodating

[1811] The server performs multilingual text translation to generate subtitles and narration in the user's language, and customizes the content based on the user's culture and customs. By using translation software such as Google Translate API, it can accommodate users from different cultures.

[1812] 6. Privacy and Data Security

[1813] The server protects uploaded digital data using the latest encryption technology. AI models also anonymize data and mask unnecessary parts to maximize user privacy. Security software such as AES (Advanced Encryption Standard) is used.

[1814] 7. Generate and preview the movie

[1815] The server combines all the elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. Users can check the preview and provide feedback through the interface on any necessary corrections (e.g., changing the order of scenes or replacing music).

[1816] 8. Final output and sharing

[1817] The server generates the final version of the movie incorporating the user's feedback and corrections. The user can choose the export format of the movie file (e.g., high-resolution video file or streaming link) according to their preference. The user receives the final movie file and can save it locally or share it on an online platform.

[1818] As a concrete example, consider a scenario where a user uploads photos and videos of clothes they tried on in a virtual store, and a shopping experience video is generated based on those photos and videos. Here is an example of a prompt to be input to the generative AI model:

[1819] "Create an inspiring shopping experience video using photos and videos of users trying on different outfits in a virtual store. At the end of the video, highlight the user's favorite outfit and play inspiring music in the background."

[1820] This prompt allows the generative AI model to generate a user experience movie that can then be previewed and shared.

[1821] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1822] Step 1:

[1823] Users upload photos, videos, and audio files through dedicated applications or websites. The input is the user's digital data, and the output is data stored on the server. Specifically, the user selects the captured data and presses the upload button, which sends the file to the server.

[1824] Step 2:

[1825] The server analyzes the uploaded data and classifies photos, videos, and audio recording files into their respective categories. The input is the uploaded digital data, and the output is the classified data. Specifically, image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recording files are also converted into text using speech recognition technology, and important messages and keywords are extracted. This processing is done using OpenCV and TensorFlow.

[1826] Step 3:

[1827] The server references the user's specified theme and wishes, and automatically generates an optimal story based on key moments and messages. The input is classified data and the user's theme settings, and the output is a storyboard. Specifically, a generative AI model analyzes the data and constructs the story flow. This process generates a coherent and moving story.

[1828] Step 4:

[1829] Based on the story, the server selects music from a library to enhance emotions and adds appropriate effects to each scene. The input is a storyboard, and the output is the final video scene. Specifically, PIL (Python Imaging Library) and moviepy are used to insert music and effects into the video. This process creates visually and aurally appealing content.

[1830] Step 5:

[1831] The server performs multilingual text translation to generate subtitles and narration in the user's specified language. It also customizes the content based on the specified culture and customs. The input is the story text information and the user's specified language, and the output is the translated subtitles and narration. Specifically, it uses the Google Translate API or similar to translate the text data and generate the required subtitles and narration.

[1832] Step 6:

[1833] The server protects uploaded digital data using the latest encryption technology. Furthermore, the AI ​​model anonymizes the data and masks unnecessary parts to ensure maximum user privacy. The input is the user's digital data, and the output is protected data. Specifically, the data is encrypted using security technologies such as AES (Advanced Encryption Standard).

[1834] Step 7:

[1835] The server aggregates all elements and generates the final movie file. The generated movie is temporarily saved and made available for users to preview. The input is the aggregated content, and the output is the generated movie file. Specifically, software such as moviepy is used to combine all scenes and effects and create the final movie file.

[1836] Step 8:

[1837] The user previews the generated movie and provides feedback through the interface on any necessary modifications (e.g., changing the order of scenes or replacing music). The input is the generated movie file and the user's feedback, and the output is the modified movie file. Specifically, the server performs the modifications based on the feedback provided by the user.

[1838] Step 9:

[1839] The server generates the final version of the movie incorporating the modifications based on the user's feedback. The export format (e.g., high-resolution video file or streaming link) can be selected according to the user's preference. The input is the modified movie file, and the output is the final movie file. Specifically, we provide the function to export the movie file in the format desired by the user.

[1840] Step 10:

[1841] Users receive the final movie file and can save it locally or share it on an online platform. The input is the final movie file, and the output is the shared movie. Specifically, the generated movie file can be uploaded to social media or cloud storage to be shared with other users.

[1842] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1843] This invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by the user. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to generate more emotionally rich stories. To implement the system, it is necessary to design and implement programs based on the following means.

[1844] 1. Data collection

[1845] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. These data are temporarily stored on the device.

[1846] 2. Data transmission

[1847] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[1848] 3. Data Analysis and Classification

[1849] Server: Receives uploaded data and classifies photos, videos, and audio recordings. Image recognition technology is used to identify scenes in photos and videos and tag specific moments. Audio recordings are also converted to text using speech recognition technology to extract important messages and keywords.

[1850] 4. Emotional Recognition

[1851] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[1852] 5. Build a story

[1853] Server: Based on the user's specified themes and wishes, and taking into account the emotional information recognized by the emotion engine, the server automatically generates a moving story. The AI ​​model determines the appropriate scene order and narrative to build a compelling story.

[1854] 6. Music and Effects Selection

[1855] Server: Based on the story, select music from the library to enhance the emotion and add effects to suit the scene, such as fade in / out and slow motion.

[1856] 7. Multilingual and culturally adaptable

[1857] Server: Generates subtitles and narration based on the user-specified language. Uses text translation tools for multilingual support and adjusts content for cultural adaptation.

[1858] 8. Privacy and Data Security

[1859] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1860] 9. Generate and preview the movie

[1861] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and users can view it using the preview link.

[1862] Users: can use the preview link to see the generated movie and provide feedback, such as requests for changes to the scene order or music.

[1863] 10. Receiving and Responding to Feedback

[1864] Server: Receives user feedback and modifies the movie based on the instructions. The modified movie is then regenerated and saved.

[1865] 11. Final output and sharing

[1866] Server: Generates the final version of the movie and provides it to the user in the export format of their choice, such as a high-resolution video file or a streaming link.

[1867] Device: Receive the final movie file and save it locally or share it on an online platform.

[1868] Specific examples

[1869] For example, consider the case of creating a wedding movie.

[1870] 1. Data collection

[1871] Users: Upload photos and videos from the wedding day, as well as recorded messages from friends and family.

[1872] 2. Data transmission

[1873] Terminal: Sends these uploaded data to the server.

[1874] 3. Data Analysis and Classification

[1875] Server: Recognizes and tags important moments in photos and videos (e.g., the promise kiss, the ring exchange), and extracts inspirational messages from audio recordings.

[1876] 4. Emotional Recognition

[1877] Server: Analyzes facial expressions and tone of voice from video and audio to recognize the emotions of the bride and groom and guests. For example, it identifies touching moments by detecting smiles and tears.

[1878] 5. Build a story

[1879] Server: Based on the extracted scenes and messages and the emotional information recognized by the emotion engine, a moving story of the entire wedding is constructed.

[1880] 6. Music and Effects Selection

[1881] Server: Add inspiring music and apply effects appropriate to the scene, including fade in / out, slow motion, etc.

[1882] 7. Multilingual and culturally adaptable

[1883] Server: For international weddings, generate subtitles and narration translated into each country's language.

[1884] 8. Privacy and Data Security

[1885] Server: All data is encrypted and secured.

[1886] 9. Generate and preview the movie

[1887] User: Preview the generated movie and request corrections if necessary.

[1888] 10. Receiving and Responding to Feedback

[1889] Server: Based on user feedback, the movie is regenerated and modifications are made.

[1890] 11. Final output and sharing

[1891] Server: Generates the final version of the movie and serves it to the user.

[1892] Device: Download movies and finally save and share them.

[1893] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[1894] The processing flow will be explained below.

[1895] Step 1:

[1896] User: Uploads photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. Uploaded data is temporarily stored on the device.

[1897] Step 2:

[1898] Terminal: The data uploaded by the user is sent to the server via the Internet. The data also includes metadata (e.g., date, time, location, etc.).

[1899] Step 3:

[1900] Server: Receives the uploaded data and sorts the photos, videos, and audio recording files into their respective categories. This process also includes format checking and conversion of the data.

[1901] Step 4:

[1902] Server: Analyzes photo data using image recognition technology, automatically identifying faces and tagging distinctive scenes (e.g., smiles or touching moments).

[1903] Step 5:

[1904] Server: Analyzes video data using the video analysis module, detects scene changes, and extracts important moments (e.g., the engagement kiss or ring exchange).

[1905] Step 6:

[1906] Server: Using speech recognition technology, the audio recordings are converted into text, and then important messages and keywords are extracted and classified.

[1907] Step 7:

[1908] Server: The emotion engine is used to recognize the user's emotions from the uploaded video and audio. For example, it analyzes facial expressions and tone of voice to identify moments when the user is particularly moved or joyful.

[1909] Step 8:

[1910] Server: Automatically generates an emotional storyboard based on the user's specified themes and wishes, and also takes into account recognized emotional information. The AI ​​model determines the order of scenes and the progression of the story, building a coherent story.

[1911] Step 9:

[1912] Server: Based on the story, select music from the library to enhance the emotions and add effects to each scene, such as fade in / out and slow motion.

[1913] Step 10:

[1914] Server: Generates subtitles and narration based on the user's language preference. For multilingual support, it uses text translation tools and adjusts the content to take cultural adaptation into account.

[1915] Step 11:

[1916] Server: All uploaded data is encrypted using the latest encryption technology to ensure privacy and data security.

[1917] Step 12:

[1918] Server: Integrates all elements and generates the first movie file. The generated movie is temporarily saved and a preview link is provided to users.

[1919] Step 13:

[1920] Users: Use the preview link to view the generated movie and submit correction requests through the interface, such as changing the order of scenes or the music.

[1921] Step 14:

[1922] Server: Receives user feedback and modifies the movie based on the user's instructions. The modified movie is regenerated and temporarily stored.

[1923] Step 15:

[1924] Server: Generates the final movie file and exports it in the format of your choice (e.g. high-resolution video file or streaming link).

[1925] Step 16:

[1926] Terminal: Receive the final movie file and save it locally. Furthermore, if you want to share it on an online platform, you can easily do so using a dedicated application.

[1927] In this way, the system of the present invention utilizes an emotion engine to automatically generate even more moving movies based on digital data that records a user's special moments, and provides them as personalized keepsakes.

[1928] Example 2

[1929] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1930] In recent years, the widespread use of smartphones and digital cameras has increased the opportunities for users to save various events and everyday moments as digital data (photos, videos, and audio recording files). However, editing this data to create moving and personalized movies requires advanced editing skills and time, making it difficult for average users. Conventional methods have been problematic in that the editing process is cumbersome and it is difficult to automatically generate emotionally appealing stories, leaving users unable to obtain satisfactory results.

[1931] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to upload photos, videos, and audio recording files; a means for transmitting the uploaded data to the server via the Internet; a means for the server to analyze the received data, classify the content, and extract important scenes and messages; a means for analyzing the user's emotions using emotion recognition technology; a means for automatically generating a story based on the user's preferences and themes, taking into account the emotional information; a means for adding music and effects to the story; a means for multilingual support and cultural adaptation; a means for presenting the generated movie to the user and making corrections based on the user's feedback; a means for ensuring privacy protection and data security; and a means for exporting and sharing the final movie file in a specified format. This allows even ordinary users to easily automatically generate moving and personalized movies and obtain high-quality results.

[1932] "User" refers to any individual or organization that uses this system.

[1933] "Photo" refers to still image data that a user takes using a digital device and uploads to the system.

[1934] "Video" refers to moving image data that a user shoots using a digital device and uploads to the system.

[1935] "Audio recording file" refers to a file that contains audio data recorded by a user or another person.

[1936] "Upload" refers to a user sending a photo, video, or audio recording file to the system through a dedicated application or website.

[1937] "Internet" refers to a global network for data communication.

[1938] "Server" refers to a computer system that receives, analyzes, sorts, and processes data sent by users.

[1939] "Analysis" refers to the process by which the server understands and interprets the content of photographs, videos, and audio recordings based on data.

[1940] "Classification" refers to the process by which the server divides the data it receives into specific categories.

[1941] "Important scenes and messages" refer to notable moments or words that the user finds interesting or that the system automatically selects.

[1942] "Emotion recognition technology" refers to the technology that a system uses to analyze and identify a user's emotions from images and audio.

[1943] "Automatic story generation" refers to the process in which a system constructs a series of events into a narrative based on the user's wishes, themes, and emotional information.

[1944] "Music and Effects" refers to background music and visual effects added to enhance the visual and auditory impression of the movie.

[1945] "Multilingual" refers to the ability to display content or provide narration in different languages.

[1946] "Cultural adaptation" refers to adjusting content to suit the user's culture.

[1947] "Feedback" refers to the act of a user conveying to the system their opinions and requests for corrections and improvements to a movie.

[1948] "Privacy protection" refers to protecting users' personal information and data from being leaked to third parties.

[1949] "Data security" refers to measures taken to keep data protected from unauthorized access and tampering.

[1950] "Export" refers to the act of saving or outputting the final movie file in a particular format.

[1951] "Sharing" refers to sharing the generated movie file or link with other users and platforms.

[1952] The present invention is a system that automatically generates moving and personalized movies based on digital data (photos, videos, and audio recording files) provided by users. Specific procedures for implementing the present invention and the hardware and software used are described below.

[1953] This system is primarily composed of users, devices, and a server. Users upload data using a dedicated application or website, and their devices send the data to the server via the Internet. The server analyzes the data, generates a story, and provides the final movie file to the user. The specific operation is explained below.

[1954] Data collection

[1955] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website, and these data are temporarily stored on the device.

[1956] Sending data

[1957] The device sends the data uploaded by the user to a server via the Internet, including metadata (e.g., date, time, location, etc.).

[1958] Data analysis and classification

[1959] The server analyzes the received data and classifies each photo, video, and audio recording file using image recognition technologies (e.g., Google Cloud Vision API) and speech recognition technologies (e.g., Google Cloud Speech-to-Text). Scenes in the photos and videos are identified and specific moments (e.g., the engagement kiss or the ring exchange) are tagged. Audio recordings are converted to text and key messages and keywords are extracted.

[1960] Emotion recognition

[1961] The server uses emotion recognition technology (e.g., Microsoft Azure Emotion API) to recognize the user's emotions from video and audio, analyzing smiles, tears, tone of voice, etc. to identify emotional moments.

[1962] Building a story

[1963] The server automatically generates a moving story based on the user's preferences and themes, taking into account emotional information. This story generation uses an AI model (e.g., OpenAI GPT) to determine the order of movie scenes and the content of the narration.

[1964] Music and effects selection

[1965] Based on the story, the server selects music from a library (e.g., Epidemic Sound) to enhance the emotions and adds effects (fade in / out, slow motion, etc.) that match the scene.

[1966] Multilingual and culturally accommodating

[1967] The server generates subtitles and narration based on the language specified by the user, and uses text translation tools (e.g., Google Translate API) to support multiple languages, adjusting the content for cultural adaptation.

[1968] Privacy and Data Security

[1969] The server uses the latest encryption technology (e.g., AES-256) for all uploaded data to ensure privacy and data security. Data is sent and received using SSL / TLS encryption, and data is encrypted when stored.

[1970] Generate and preview the movie

[1971] The server combines all the elements to generate a first version of the movie file, stores it temporarily, and makes it available to users via a preview link.

[1972] Users can use the preview link to view the resulting movie and submit requests for changes to the scene order or music, if necessary.

[1973] Receiving and responding to feedback

[1974] The server receives feedback from the user and modifies the movie based on the instructions, and the modified movie is regenerated and saved.

[1975] Final output and sharing

[1976] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link.

[1977] The device will receive the final movie file and store it locally, or you can share it on an online platform.

[1978] Examples of prompt statements

[1979] Below are some example prompts to input to the generative AI model:

[1980] Create a moving wedding video using photos and videos from your wedding, along with messages from friends and family. Highlight the bride and groom's smiling and tearful moments, and add music appropriate to their special day. For an international wedding, add subtitles in English and Japanese to reflect cultural elements. Specifically, use emotional music for the groom's entrance and select a song that gradually builds up during the bride's entrance. Additionally, highlight the vows and ring exchange in slow motion, and add moving guest messages.

[1981] Through the above steps, the present invention utilizes an emotion engine to automatically generate an even more moving movie based on digital data that records a user's special moments, and provides it as a personalized keepsake.

[1982] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1983] Step 1: Collect data

[1984] Users upload photos, videos, and audio recordings related to events such as weddings and birthdays through a dedicated application or website. The input is the user's own digital data. The uploaded data is temporarily stored on the device and becomes the output. For example, a user can drag and drop photos from their wedding day or audio files containing messages from friends into the application, and then enter the event theme and desired details into a form.

[1985] Step 2: Sending data

[1986] The device sends the data uploaded by the user to the server via the Internet. The input includes data temporarily stored on the device and metadata (date, time, location, etc.). The output is data sent to the server and received by the server. Specifically, the device securely sends this data to the server using the HTTPS protocol.

[1987] Step 3: Analyze and classify the data

[1988] The server analyzes and classifies the data it receives. The input is the data received by the server. The output is photos, videos, and audio recordings sorted into their respective categories, with each file tagged with important scenes and keywords. For example, the server uses Google Cloud Vision API to identify scenes in photos and videos and tag specific moments, such as the bride and groom's kiss and ring exchange. It also uses Google Cloud Speech-to-Text to convert audio recordings into text and extract important messages and keywords.

[1989] Step 4: Recognize emotions

[1990] The server uses emotion recognition technology to analyze user emotions from video and audio. The input is analyzed and classified video and audio data. The output is an emotional tag attached to each piece of data. Specifically, it uses the Microsoft Azure Emotion API to analyze smiles, tears, tone of voice, and other emotions to identify moments when the bride and groom and guests are emotional.

[1991] Step 5: Build your story

[1992] The server automatically generates a story based on the user's wishes and themes, taking into account emotional information. The inputs are data tagged with emotions and the wishes and themes provided by the user. The output is a movie plan that determines the order of scenes and narration for a moving story. Specifically, it uses OpenAI GPT to construct the story's narrative and determine the placement of scenes.

[1993] Step 6: Choose music and effects

[1994] The server selects music and effects based on the story. The input is a story plan. The output is a story plan with music and effects applied to scenes to enhance the emotions. Specific operations include selecting appropriate music from the Epidemic Sound library and applying effects such as fade-in / out and slow motion to scenes.

[1995] Step 7: Multilingualism and cultural adaptation

[1996] The server generates subtitles and narration based on the language specified by the user. The inputs are the language information specified by the user and a story plan. The output is a story plan that includes subtitles and narration in multiple languages. Specifically, the system uses the Google Translate API to translate the text into multiple languages ​​and edits the content to take cultural adaptation into account.

[1997] Step 8: Privacy and Data Security

[1998] The server uses encryption technology when sending, receiving, and storing data. The inputs include user data stored on the server and data being sent and received. The output is encrypted data. Specifically, data security is ensured using AES-256 data encryption and SSL / TLS protocols.

[1999] Step 9: Generate and preview your movie

[2000] The server integrates all elements to generate a first-run movie file and temporarily stores it. The input is the final story plan. The output is the first-run movie file. Specifically, video editing software is used to generate the movie based on the pre-planned layout.

[2001] The user uses the preview link to view the generated movie. The input is the generated movie file. The output is feedback based on the preview results.

[2002] Step 10: Receiving and responding to feedback

[2003] The server receives user feedback and modifies the movie based on the instructions. The inputs are the user's feedback and the original movie file. The output is a modified movie file that reflects the feedback. Specifically, the server reorders the requested scenes and adjusts the music to generate a new movie.

[2004] Step 11: Final output and sharing

[2005] The server generates the final version of the movie and delivers it to the user in the format of their choice, such as a high-resolution video file or a streaming link. The input is the final, modified movie plan. The output is the final movie file. Specific operations include generating an HD MP4 file or a YouTube streaming link as the final output.

[2006] The device receives the final movie file and stores it locally. The input is the final movie file provided by the server. The output is a locally stored movie file, which the user can also upload to an online platform for sharing.

[2007] (Application example 2)

[2008] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2009] Conventional digital content generation systems that use photos, videos, and audio files have struggled to automatically generate emotionally rich stories. Furthermore, there are challenges, such as generating movies that effectively reflect the user's emotions, supporting multiple languages, and protecting privacy. In particular, there is a need for a system that can visually and emotionally recreate users' memories and share them in high resolution with simple smartphone application operations.

[2010] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading photos, videos, and audio recording files, means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages, and means for automatically generating an emotional story based on the user's wishes and themes. This makes it possible to automatically generate an emotional digital story that effectively reflects the user's emotions using an emotion engine.

[2011] A "photograph" is a file that digitally records a still image.

[2012] A "video" is a file that contains a digital recording of moving images.

[2013] An "audio recording file" is a file that records audio in digital format.

[2014] "Uploading" means sending local data owned by the user to a server via the Internet.

[2015] "Data analysis" refers to the use of machine learning and algorithms to process uploaded digital content and extract specific information.

[2016] "Content classification" refers to grouping photos, videos, and audio recording files based on their format and characteristics.

[2017] "Key Scene and Message Extraction" refers to automatically identifying moments and meaningful messages within digital content that deserve special emphasis.

[2018] "User's wishes and themes" are requests and objectives for a story or movie specified by the user.

[2019] "Automatic generation of emotional stories" is the process of automatically constructing emotional stories based on extracted scenes and messages.

[2020] ...

Claims

1. A means to upload photos, videos, and audio recordings; A means for analyzing the uploaded data, classifying the content, and extracting important scenes and messages; A means to automatically generate moving stories based on user wishes and themes, A way to add music and effects to your moving story, Multilingual and cultural adaptation measures; a means for presenting the generated movie to a user and making modifications based on the user's feedback; measures to ensure privacy and data security; A system that includes the means to export and share the final movie file in the format of your choice.

2. 10. The system of claim 1, further comprising means for analyzing and categorizing user uploaded data by format, facial recognition technology and key moment tagging.

3. 10. The system of claim 1, further comprising means for converting a recorded conversation file into text using voice recognition technology and extracting important messages and keywords.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A