system

The system automates the creation and publication of high-quality video content by capturing, compressing, analyzing, and editing user-generated videos, addressing the complexity of scene extraction and personal information protection.

JP2026071044APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Creating high-quality video content requires specialized techniques and labor, and extracting important scenes from individual videos while protecting personal information is complex and difficult to automate.

Method used

A system that includes a recording device to capture video based on user-specified conditions, a compression method to convert video into a predetermined format, a transmission device to upload data to a central processing unit, an analysis device to analyze video data using AI technology, a generation device to edit the data, and a posting device to publish the content automatically on online platforms, allowing users to review and adjust the content through a user interface.

Benefits of technology

Enables users to generate and publish high-quality video content without specialized technical skills, efficiently extracting important scenes and protecting personal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071044000001_ABST
    Figure 2026071044000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A system comprising: a recording device that automatically captures video based on conditions; a compression device that converts the captured video data into a predetermined format; a transmission device that uploads the compressed video data to a central processing unit via a network; an analysis device that analyzes the video data using AI technology and extracts important scenes; a generation device that edits the extracted video data using AI technology and generates a visual medium; a user interface that presents the generated visual medium to the user and accepts adjustments; and a posting device that publishes the adjusted visual medium to an online platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Creating a video blog requires specialized techniques and labor related to shooting and editing, which is a high hurdle for many users. Additionally, there is also an issue that it is difficult for users without shooting or editing skills to generate high-quality video content by themselves. Furthermore, there is a problem that extracting important scenes from individual videos and editing considering the protection of personal information is complex and difficult to automate.

Means for Solving the Problems

[0005] This invention efficiently transfers data by having a recording device automatically capture video according to user-specified conditions and convert the video into a predetermined format using a compression means. Subsequently, a transmission device uploads the data to a central processing unit (server) via a network, and an analysis device analyzes the video data using AI technology to extract important scenes. A generation device further edits the data extracted by AI technology and generates a visual medium. Here, the user can review the generated visual medium through a user interface and easily make adjustments. Finally, a posting device automatically publishes this adjusted visual medium on an online platform, thereby solving the above problem.

[0006] A "recording device" is a hardware device that has the function of automatically capturing video based on conditions specified by the user.

[0007] A "compression method" is a technology that has the function of converting captured video data into a predetermined compression format and reducing the data size.

[0008] A "transmitting device" is hardware or software that has the function of transmitting compressed video data to a central processing unit via a network.

[0009] A "central processing unit" is a centralized computing system on a platform that processes data received from a transmitting device and manages its analysis and editing.

[0010] An "analysis device" is software or hardware that uses AI technology to analyze video data and extract important scenes and information.

[0011] A "generation device" is a device that executes the process of generating visual media using AI technology based on analyzed data.

[0012] A "user interface" refers to the environment or device that allows a user to view and adjust the generated visual media.

[0013] A "posting device" is a device or platform that has the functionality to publish a prepared visual medium onto an online platform.

[0014] "Visual media" refers to the final video content generated through AI editing, and takes a form that users can view as visual information. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is configured as a system that automatically records and edits a user's daily activities. By wearing a small recording device, the user can easily record various scenes of their daily life. The recording device detects the user's movements and location information and automatically activates the camera based on set conditions. As a result, the user does not need to consciously operate the camera, and can capture video in a natural way.

[0037] The captured video data is effectively compressed within the recording device using compression methods and then transferred to a server via the network. The server receives this data and performs analysis using AI technology with an analysis device. In particular, it is possible to extract specific actions or important scenes from the video and omit unnecessary parts.

[0038] Subsequently, the generation device edits the video based on the analysis results. During the editing process, appropriate cuts and scene order are automatically adjusted, and mosaic processing and text addition to the audio are performed. Here, AI-powered natural language processing is used to automatically generate text content and embed it within the video.

[0039] The generated visual media is provided to the user through a user interface. Here, the user can preview the content and make simple adjustments. This adjustment process is designed with the user experience in mind, and features an intuitive interface.

[0040] The final visual medium is automatically published on online platforms via the posting device. Users can seamlessly share content across multiple social networking platforms with a single operation. This entire process allows users to automatically generate and publish high-quality video blogs without relying on specialized technical skills. A concrete example is a scenario where a recording device captures scenes while a user visits a tourist destination during a trip, and the edited video is immediately published online upon their return home.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device uses sensors to detect the user's movements and location, and when pre-set conditions are met, it automatically activates the camera and starts capturing video.

[0044] Step 2:

[0045] The device converts the captured video into a predetermined compression format (e.g., H.264) and temporarily stores it in its internal storage. At the same time, metadata related to the capture (such as the date and time of capture and location information) is also added.

[0046] Step 3:

[0047] When the terminal detects a stable network connection, it automatically begins sending compressed video data stored on the recording device to the server. At this time, it uploads the data using the most suitable method, either Wi-Fi or mobile data.

[0048] Step 4:

[0049] The server receives the transmitted video data and performs analysis using AI technology through an analysis device. This analysis includes detecting important scenes, voice recognition, and detecting specific objects in the video.

[0050] Step 5:

[0051] The server uses a generation device to edit the video based on the analysis results. During the editing process, the system automatically cuts and rearranges scenes, applies mosaic effects, adds background music, and generates and places subtitles as needed.

[0052] Step 6:

[0053] The server provides the user with a preview of the edited visual media through the user interface. The user can review this preview and send feedback and minor adjustments to the server through the interface.

[0054] Step 7:

[0055] The server makes the necessary changes based on user feedback and then exports the final visual medium as the completed version.

[0056] Step 8:

[0057] The server automatically publishes the final visual version on the designated online platform using a posting device, allowing users to share the content via social media.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] The process of recording daily life and later editing and sharing that footage in high quality is time-consuming and requires specialized skills, making it a significant burden for the average user. Therefore, there is a need for a user-friendly, intuitive, and automated video editing and sharing system.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for a recording device to automatically acquire video based on conditions, means for a compression device to convert the acquired video information into a predetermined format, and means for an analysis device to analyze the video information using artificial intelligence technology and select specific scenes. This makes it possible to automatically extract important scenes from daily video recordings and efficiently edit and transmit them.

[0063] A "recording device" refers to a device that automatically acquires video footage of a user's daily activities based on specified conditions.

[0064] A "compression device" refers to a device that converts acquired video information into a predetermined format so that it can be efficiently stored or transmitted.

[0065] A "transmitting device" refers to a device that transmits compressed video information to a central control unit via a communication network.

[0066] An "analysis device" refers to a device that uses artificial intelligence technology to analyze video information and select specific scenes.

[0067] A "generation device" refers to a device that edits selected video information based on analysis to create visual media.

[0068] A "user interface" refers to a screen or control system that presents generated visual information to the user and allows them to make corrections using an intuitive method.

[0069] A "posting device" refers to a device used to transmit modified visual information media to an electronic platform.

[0070] "Visual information media" refers to video content produced through editing.

[0071] "Artificial intelligence technology" is a general term for computer science methods that understand data, learn patterns, or make predictions.

[0072] The system of this invention begins with a recording device worn by the user. The user carries a small recording device that incorporates a camera and sensors. This recording device, with pre-installed software, monitors the user's activities in real time and automatically acquires video when set conditions are met. For example, the camera is activated when a specific geographical location is reached using a position sensor and clock function. This process utilizes conventional digital camera technology and motion detection algorithms.

[0073] The acquired video information is compressed through a compression device within the recording device. Standardized video codec technologies such as H.264 and HEVC are used for this compression. The compressed data is then transferred to a server using wireless communication. Wi-Fi and 4G / 5G networks are utilized for this operation.

[0074] The server receives compressed video data transmitted over the network and analyzes its contents using an analysis device. This analysis utilizes artificial intelligence technology, particularly deep learning models, to extract specific scenes and actions from the video. Specific software examples include frameworks such as TENSORFLOW® and PyTorch. For example, it's possible to extract only scenes related to tourist destinations or attractions from video footage taken by a user during a trip.

[0075] Based on the analysis results, the generation device edits the video. During the editing process, the sequence is optimally rearranged, unnecessary parts are deleted, and mosaic processing is performed to prevent the identification of individuals. In addition, AI-based natural language processing technology is used to add subtitles that involve converting audio data into text. For example, Google® Cloud Speech-to-Text, a speech recognition software, is used to convert audio into text.

[0076] The user interface provides the user with the generated visual medium, allowing for preview and fine-tuning manually. This interface is optimized for touch operation, ensuring intuitive and easy usability. For example, users can tap scenes on a tablet or smartphone and drag and drop to change the sequence position.

[0077] Finally, the modified visual media is automatically published to online platforms via the server's posting system. Users can seamlessly share the video to multiple social networking platforms with just a few clicks.

[0078] As an example of a prompt, the generative AI model can receive instructions for the system in the form of, "Please describe the process of automatically recording visits to tourist spots during a trip and then automatically editing the video and posting it on social media after returning home."

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user wears a small recording device and goes about their daily activities. This device uses a built-in camera and sensors to detect the user's current location and movements. When pre-set conditions (e.g., a specific location or time) are met, the camera automatically activates and video is captured. The input is the user's movement data and current location, and the output is the raw video data captured when the conditions are met.

[0082] Step 2:

[0083] A compression device built into the recording device compresses the acquired video data. Specifically, it reduces the amount of data using compression technologies such as H.264 and HEVC. The input is raw video data, and the output is a compressed video file. This compression enables efficient data transfer.

[0084] Step 3:

[0085] The compressed video file is transmitted to the server via wireless communication (Wi-Fi or mobile data communication). The terminal uses the network through the transmitting device. The input is compressed video data, and the output is video data securely stored on the server.

[0086] Step 4:

[0087] The server passes the received video data to the analysis device. The analysis device uses a deep learning algorithm with a generative AI model to analyze the video data and automatically extract specific important scenes. The input is compressed video data on the server, and the output is video data with the selected scenes.

[0088] Step 5:

[0089] Based on the analysis results, the server's generation device edits the video. Specific actions include rearranging scenes, applying mosaic effects, and adding subtitles. If audio data is available, natural language processing is used to convert it to text and embed subtitles into the video. The input is the analyzed video data, and the output is the edited visual information medium.

[0090] Step 6:

[0091] The server sends the generated visual information medium to the user's terminal. The user reviews and adjusts the edited content through a dedicated user interface. For example, fine-tuning can be done using sliders or touch controls. The input is the edited visual information medium, and the output is the visual information medium with final adjustments made by the user.

[0092] Step 7:

[0093] Ultimately, the server's posting system publishes the adjusted visual information medium to the designated online platform. Users can select various social networking services and simultaneously distribute the video to multiple platforms with just a few operations. The input is the final adjusted visual information medium, and the output is the published video content.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] In modern times, there is a growing demand for individuals to record their daily activities and then distribute them as high-quality video content. However, traditional methods require a lot of manual work and advanced skills from recording to editing and distribution, making it difficult for users without specialized skills. Furthermore, extracting and editing important scenes from the recorded footage is time-consuming, hindering the smooth creation and distribution of content.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to analyze video data using AI technology and extract important scenes, and means for delivering user-generated content based on visual media generated on a smart device. This makes it possible for users to easily record, edit, and distribute their daily activities without requiring specialized knowledge.

[0099] A "recording device" is a device that has the function of automatically capturing video based on conditions set by the user.

[0100] "Compression means" refers to means that have the function of converting captured video data into a predetermined format, making it efficiently stored and transmitted.

[0101] A "transmission device" is a device used to upload compressed video data to a central processing unit via a network.

[0102] An "analysis device" is a device that uses AI technology to analyze video data and extract important scenes or specific actions.

[0103] A "generation device" is a device equipped with the function of editing analyzed video data and generating the final visual medium.

[0104] A "user interface" is an interface that presents generated visual media to the user and allows them to adjust the visual media.

[0105] A "posting device" is a device for publishing a modified visual medium on an online platform.

[0106] A "smart device" is a device that has the ability to process multimedia content and distribute it over the internet.

[0107] Regarding embodiments for carrying out the invention, this invention is a system that automatically generates and distributes high-quality video content of a user's daily activities. This system involves the coordinated operation of various components, including a recording device, an analysis device, a generation device, a user interface, and a transmission device.

[0108] First, the recording device automatically captures important scenes from daily life based on conditions set by the user. Specifically, the camera detects the user's movements and location information and starts capturing video. Smartphones and wearable devices are often used as recording devices.

[0109] The captured video data is converted to a predetermined format by a compression method, preparing it for efficient transmission. The compressed data is then uploaded by the transmission device to the server's central processing unit via the network.

[0110] The analysis system on the server uses a generative AI model to analyze video data and automatically extract important scenes and actions. It also analyzes audio data as needed and generates text using natural language processing techniques. AI libraries such as TensorFlow and OpenCV are used for this analysis process.

[0111] The analyzed data is passed to a generation device, where it is edited into a visual medium using AI technology. This process automatically performs operations such as mosaic processing and adding text to audio. The generated visual medium is provided to the user through a user interface, which can be adjusted as needed. This interface is designed for intuitive user operation.

[0112] The final visual medium is published on an online platform via a posting device. At this stage, user-generated content is delivered directly from smart devices. For example, in a video themed around a family barbecue, the camera can capture smiles and cooking scenes, and natural conversations and cheers can be transcribed and added to the video. After returning home, this video can be easily reviewed and shared on various platforms.

[0113] An example of a prompt would be: "Generate a scenario that visualizes a family enjoying a barbecue together and displays their particularly enjoyable conversation as text."

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The recording device detects the user's movements and location information in real time and determines when the trigger conditions for video capture are met. Based on this, the camera automatically activates and acquires video data. The input is user movement and location data, and the output is captured video data. Specifically, smartphones and wearable devices are used.

[0117] Step 2:

[0118] The compression mechanism on the terminal compresses the captured video data into a predetermined format. The input is the captured video data, and the output is the compressed video data. This process applies a highly efficient compression algorithm to reduce the data size.

[0119] Step 3:

[0120] The terminal's transmitting device uploads compressed video data to the central processing unit via the network. The input is compressed video data, and the output is data stored on the server. This step uses a secure protocol to ensure the reliability of data transfer.

[0121] Step 4:

[0122] The server uses an analysis device to analyze uploaded video data with AI technology and extract important scenes and actions. The input is video data stored on the server, and the output is a list of extracted important scenes. A generative AI model is used for scene recognition and action detection.

[0123] Step 5:

[0124] The server operates the generation device, edits the video based on key scenes, and generates the visual medium. The input is a list of extracted key scenes, and the output is the completed visual medium. Specifically, the order of the scenes is automatically adjusted, and mosaic processing or text addition is performed as needed. Using natural language processing technology, the server analyzes the audio data and overlays the generated text onto the video.

[0125] Step 6:

[0126] The user interface presents the generated visual medium to the user and allows for easy adjustments. The input is the completed visual medium, and the output is the visual medium adjusted by the user. The user can intuitively edit through the interface.

[0127] Step 7:

[0128] The terminal's posting device publishes the customized visual medium to an online platform. The input is the user-customized visual medium, and the output is the published online content. This step seamlessly distributes content across multiple platforms.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention provides a system that integrates a recording device, an emotion engine, an analysis device, and a generation device, enabling users to automatically generate video content while naturally incorporating their own emotions and actions. Users can easily record their daily activities by wearing a small recording device. The recording device detects the user's movements and location information and automatically captures video based on specific conditions.

[0131] The recorded video data is efficiently compressed using a compression method and then transmitted to a server via the network. The server processes the received video data with an analysis device and uses an emotion engine to recognize emotions from the user's facial expressions and voice. This makes it possible to identify important scenes based on the user's emotions within the video data.

[0132] The generation device takes into account the emotional data recognized by the emotion engine and performs appropriate video editing. Specifically, it customizes the visual medium to match the user's emotions by changing the tone of the video and automatically setting background music according to the user's feelings of joy or sadness. Because this process is fully automated by AI technology, users do not need to perform any technical operations and can use it intuitively.

[0133] The completed visual content is presented to the user through a user interface, allowing them to preview the video and make simple adjustments as needed. Once edited, the posting device automatically publishes the visual content to the online platform, enabling users to share the content with a wide audience with a single click.

[0134] As a concrete example, imagine a user attending a birthday party, and a recording device captures the event. In this case, the emotion engine detects the user's smiles and cheers, and the generator edits the video to a brighter tone based on this, adding cheerful music to the background. Ultimately, a video content optimized for emotion data is generated and can be easily shared.

[0135] The following describes the processing flow.

[0136] Step 1:

[0137] The device, worn by the user, uses sensors to detect the user's movements and surrounding sounds, and automatically begins capturing video and audio when predetermined conditions are met.

[0138] Step 2:

[0139] The device compresses the captured video and audio data into a compressed format for efficient processing and temporarily stores it in its internal storage. The time of capture and location information are also recorded as metadata.

[0140] Step 3:

[0141] The device uploads compressed video data to the server when a network connection becomes available. The appropriate connection method (Wi-Fi or mobile data) is automatically selected.

[0142] Step 4:

[0143] The server uses AI technology to analyze the received video data with an analysis device, and detects important scenes and events within the video.

[0144] Step 5:

[0145] The server utilizes an emotion engine to recognize emotions from the user's facial expressions and voice in the video. For example, emotions such as smiles, surprise, and joy are extracted.

[0146] Step 6:

[0147] The generation device edits the video based on recognized emotions and analyzed data of key scenes. The video's color tone, scene flow, and background music are automatically adjusted to match the user's emotions.

[0148] Step 7:

[0149] The server provides the edited visual media to the user as a preview through the user interface. The user can review the video content and make simple editing requests if necessary.

[0150] Step 8:

[0151] The server generates the final visual medium incorporating user feedback and automatically publishes it to the online platform via the posting device. Users can then easily share the video with their audience.

[0152] (Example 2)

[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0154] Current technology makes it difficult to efficiently record users' daily activities as video and automatically generate customized visual media that responds to their emotions based on that footage. Furthermore, there is a lack of readily available methods for users to intuitively edit videos without requiring technical knowledge and easily share them on online platforms. Additionally, processing that conserves network bandwidth through video data compression and transmission while protecting user privacy is essential.

[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0156] In this invention, the server includes means for an analysis device that uses AI technology to analyze video data and identify important scenes, including the user's emotions; means for a generation device that edits video data based on the emotion data and generates visual media; and means for compression used to enable low-bandwidth communication. This enables efficient recording of daily activities, automatic generation of customized visual media based on emotions, and user-friendly, intuitive video editing and easy content sharing.

[0157] A "recording device" is a device that has the function of automatically acquiring a user's daily actions and events as video.

[0158] "Compression means" refers to a technology that converts captured video data into a predetermined format to reduce the amount of data.

[0159] A "transmission device" is a device that has the function of uploading compressed video data to an information processing device via a communication network.

[0160] An "analysis device" is a device that uses AI technology to analyze video data and identify important scenes, particularly based on emotions.

[0161] A "generation device" is a device that edits video data based on analysis results to generate visual media that matches emotions.

[0162] A "user interface" is a means of presenting a generated visual medium to the user and providing interaction to allow for easy adjustments.

[0163] A "posting device" is a device that has the function of automatically publishing a modified visual medium to an internet platform.

[0164] A "communication network" is an infrastructure for electronically sending and receiving data.

[0165] "AI technology" refers to technologies that use artificial intelligence to perform data analysis and decision-making.

[0166] "Visual media" refers to content that is edited to match the user's emotions and presented visually.

[0167] The system of this invention is a complex system including a recording device, a compression means, a transmission device, an analysis device, a generation device, a user interface, and a posting device. Specific embodiments of each component are shown below.

[0168] Recording device

[0169] Users can wear a small recording device that allows them to naturally record their daily activities. This device uses sensors to detect the user's movements and location, and automatically acquires video under specific conditions (for example, when the user is participating in a particular event).

[0170] Compression means and transmission device

[0171] The terminal receives video data acquired from the recording device, applies a video compression algorithm such as H.264 to compress the data, and then transmits the compressed data to the server via the communication network.

[0172] analysis device

[0173] The server processes the received video data using an analysis device. This analysis utilizes AI technology, particularly deep learning, to analyze the user's facial expressions and voice, and identify their emotions. Based on the analysis results, important scenes within the video are identified.

[0174] generator

[0175] The server automatically edits the video by using the analyzed emotion data in a generation device. This process uses a generation AI model to adjust the video tone and select music according to the user's emotions.

[0176] User interface and posting device

[0177] Users view the edited visual media through the device's user interface. If minor adjustments are needed, they can be made on the interface. Finally, the content is published to an online platform using a posting device and shared with a wide audience.

[0178] As a concrete example, imagine a user attending a friend's birthday party, with the recording device capturing the event. The analysis device recognizes emotions of joy from smiles and tone of voice, and the generation device edits the video into a bright-toned image, adding upbeat background music. The finished content can then be easily shared on social media.

[0179] An example of a prompt message might be, "I want you to automatically generate a video that emphasizes a cheerful atmosphere using footage from a friend's birthday party."

[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0181] Step 1:

[0182] The user acquires video using a small recording device. The recording device uses sensors to detect the user's movements and location, and automatically captures video in specific situations. The input is the user's location information and motion sensor data, and the output is the captured raw video data. As a concrete example, the recording device records the surrounding environment when the user attends a birthday party.

[0183] Step 2:

[0184] The terminal receives raw video data from the recording device. It then applies a video compression algorithm (e.g., H.264) to the received data to reduce its size. The input is raw video data, and the output is compressed video data. Specifically, the terminal compresses the video and prepares it for network transmission.

[0185] Step 3:

[0186] The terminal transmits compressed video data to the server via a communication network. The input is the compressed video data, and the output is the status indicating that the data has been successfully uploaded to the server. Specifically, the terminal uses its internet connection to send data to the server.

[0187] Step 4:

[0188] The server processes the received compressed video data using an analysis device. During this process, AI technology is used to analyze emotions and important scenes within the video. The input is the compressed video data, and the output is the analyzed emotion data and information on important scenes. Specifically, the server analyzes smiles and sounds within the video.

[0189] Step 5:

[0190] The server edits the video based on emotional data analyzed using a generator. Tone adjustments and automatic background music settings are performed. The input is the analyzed emotional data, and the output is the edited visual medium. Specifically, the server adds bright visual effects to match the emotion of joy.

[0191] Step 6:

[0192] The user previews the edited visual medium on the user interface on their device. They can adjust the tone, volume, and other aspects of the visual medium as needed. The input is the edited visual medium, and the output is improvement suggestions based on user feedback. For example, the user can change the brightness using a slider.

[0193] Step 7:

[0194] The device uses a posting device to publish the final visual medium to the online platform. The input is the finalized visual medium, and the output is the published result on the online platform. Specifically, the device clicks the publish button and shares the content.

[0195] (Application Example 2)

[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0197] In modern times, recording and viewing daily activities on video has become commonplace. However, conventional technology lacked the means to automatically edit and optimize videos in a way that resonates with the user's emotions. Furthermore, users had to perform the editing themselves, resulting in high technical hurdles and a lack of ease of use. There is a growing need for a system that dynamically adjusts the tone of the video and background sound according to emotions, and that is intuitive for users to use.

[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0199] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to extract important scenes based on the user's emotions, and means for a generation device to automatically optimize the tone and background sound of the video using AI technology based on an emotion engine. This allows users to naturally generate and share video content customized to their emotions without requiring any technical operation.

[0200] A "recording device" is a device that has the function of automatically capturing the user's activities and surrounding environment as video based on certain conditions.

[0201] "Compression means" refers to a process that converts captured video data into a predetermined format and efficiently reduces its size.

[0202] A "transmission device" is a device equipped with the function of uploading compressed video data to a central processing unit via a network.

[0203] An "analysis device" is a device that uses AI technology to automatically extract user emotions and important scenes from video data.

[0204] An "emotion engine" refers to artificial intelligence technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0205] A "generation device" is a device that uses AI technology to edit extracted video data and generate visual media optimized according to emotions.

[0206] A "user interface" is an interface that presents generated visual media to the user and allows for previewing and adjustments.

[0207] A "posting device" is a device that has the function of publishing a modified visual medium on an online platform.

[0208] "Eye-tracking information" refers to data that detects the user's eye movements and uses that information to set importance levels.

[0209] A "generative AI model" refers to an artificial intelligence model that creates and edits new video content based on given information.

[0210] A "prompt statement" refers to an input statement used to instruct a generative AI model on how to generate content.

[0211] This invention relates to a system that automatically captures user behavior and emotions and generates optimized video content. The system consists of a recording device, an analysis device, a generation device, a user interface, and a posting device.

[0212] The recording device has the function of automatically capturing video based on conditions and is implemented in the form of a small portable device or smart glasses. For example, when a user is wearing smart glasses, the device continuously measures the surrounding video and audio and captures data as needed.

[0213] The analysis device operates on a server and processes captured video data using AI technology. This device analyzes the user's emotions from facial expressions and voice, and extracts important scenes. Emotion recognition uses Microsoft® Azure® Face API and Google Cloud Vision API. For example, in a scene where a user is chatting with a friend in a cafe, the device detects a smile and marks it as a high-priority moment.

[0214] The generation device edits visual media based on emotion analysis results within the server. It utilizes a generation AI model to customize the tone and background sound of extracted scenes based on emotions. The AI ​​model operates according to pre-configured prompts. For example, in a scene where a user is enjoying a walk in a park, it would process based on the prompt: "Edit the video of the scene where the user is enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a brighter tone to the scenery."

[0215] The user interface presents the generated visual media to the user, allowing for previews and simple adjustments. This operation takes place on the user's mobile device, providing an intuitive UI. Users can visually review the generated video and switch scenes or adjust effects as needed.

[0216] Finally, the posting device automatically publishes the edited visual medium to the online platform. This feature allows users to widely share video content with a single click. For example, users can easily post a video compilation of travel memories to social media.

[0217] This system provides users with an environment where they can generate and widely share emotionally-based, customized video content without requiring any technical operation.

[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0219] Step 1:

[0220] When a user puts on smart glasses and begins an activity, the recording device starts capturing video and audio data. The input is real-time video and audio from the user's perspective, and the output is the captured raw data. The recording device selectively records data at specific events or timings based on conditions set by the user.

[0221] Step 2:

[0222] The server receives raw data transmitted from the recording device and converts it into a predetermined format using a compression method. The input is the captured raw data, and the output is compressed data with optimized capacity. This process involves data compression to eliminate unnecessary data and reduce network load.

[0223] Step 3:

[0224] The server passes compressed data to the analysis device, which uses AI technology to perform emotion analysis. The input is compressed video and audio data, and the output is emotion data estimated from the user's facial expressions and voice, as well as the extraction of important scenes that should be emphasized based on those emotions. The analysis device uses the Microsoft Azure Face API to analyze facial expressions and determine the user's emotional state. Specifically, it detects the user's smiles and expressions of surprise and identifies scenes to highlight.

[0225] Step 4:

[0226] The generation device performs video editing using a generation AI model, taking into account the emotion data output from the previous step. The input consists of key scenes and emotion data, and the output is the affected visual medium. In this process, based on prompt statements, the device automatically adjusts the tone of the video and sets background music according to instructions such as, "Edit the video of a scene of someone enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a bright tone to the scenery."

[0227] Step 5:

[0228] The terminal presents the generated visual medium to the user through a user interface and accepts simple adjustments. The input is the generated visual medium, and the output is the medium adjusted based on the user's feedback. The terminal displays the video as a preview and provides an interface that allows the user to change scenes and fine-tune effects as needed.

[0229] Step 6:

[0230] The server automatically publishes the finalized visual media to the online platform via a posting device. The input is the adjusted visual media, and the output is the published video content. In this process, the media is uploaded to a pre-configured platform account, allowing users to easily share the content through social media and other means.

[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0234] [Second Embodiment]

[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0247] This invention is configured as a system that automatically records and edits a user's daily activities. By wearing a small recording device, the user can easily record various scenes of their daily life. The recording device detects the user's movements and location information and automatically activates the camera based on set conditions. As a result, the user does not need to consciously operate the camera, and can capture video in a natural way.

[0248] The captured video data is effectively compressed within the recording device using compression methods and then transferred to a server via the network. The server receives this data and performs analysis using AI technology with an analysis device. In particular, it is possible to extract specific actions or important scenes from the video and omit unnecessary parts.

[0249] Subsequently, the generation device edits the video based on the analysis results. During the editing process, appropriate cuts and scene order are automatically adjusted, and mosaic processing and text addition to the audio are performed. Here, AI-powered natural language processing is used to automatically generate text content and embed it within the video.

[0250] The generated visual media is provided to the user through a user interface. Here, the user can preview the content and make simple adjustments. This adjustment process is designed with the user experience in mind, and features an intuitive interface.

[0251] The final visual medium is automatically published on online platforms via the posting device. Users can seamlessly share content across multiple social networking platforms with a single operation. This entire process allows users to automatically generate and publish high-quality video blogs without relying on specialized technical skills. A concrete example is a scenario where a recording device captures scenes while a user visits a tourist destination during a trip, and the edited video is immediately published online upon their return home.

[0252] The following describes the processing flow.

[0253] Step 1:

[0254] The device uses sensors to detect the user's movements and location, and when pre-set conditions are met, it automatically activates the camera and starts capturing video.

[0255] Step 2:

[0256] The device converts the captured video into a predetermined compression format (e.g., H.264) and temporarily stores it in its internal storage. At the same time, metadata related to the capture (such as the date and time of capture and location information) is also added.

[0257] Step 3:

[0258] When the terminal detects a stable network connection, it automatically begins sending compressed video data stored on the recording device to the server. At this time, it uploads the data using the most suitable method, either Wi-Fi or mobile data.

[0259] Step 4:

[0260] The server receives the transmitted video data and performs analysis using AI technology through an analysis device. This analysis includes detecting important scenes, voice recognition, and detecting specific objects in the video.

[0261] Step 5:

[0262] The server uses a generation device to edit the video based on the analysis results. During the editing process, the system automatically cuts and rearranges scenes, applies mosaic effects, adds background music, and generates and places subtitles as needed.

[0263] Step 6:

[0264] The server provides the user with a preview of the edited visual media through the user interface. The user can review this preview and send feedback and minor adjustments to the server through the interface.

[0265] Step 7:

[0266] The server makes the necessary changes based on user feedback and then exports the final visual medium as the completed version.

[0267] Step 8:

[0268] The server automatically publishes the final visual version on the designated online platform using a posting device, allowing users to share the content via social media.

[0269] (Example 1)

[0270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0271] The process of recording daily life and later editing and sharing that footage in high quality is time-consuming and requires specialized skills, making it a significant burden for the average user. Therefore, there is a need for a user-friendly, intuitive, and automated video editing and sharing system.

[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0273] In this invention, the server includes means for a recording device to automatically acquire video based on conditions, means for a compression device to convert the acquired video information into a predetermined format, and means for an analysis device to analyze the video information using artificial intelligence technology and select specific scenes. This makes it possible to automatically extract important scenes from daily video recordings and efficiently edit and transmit them.

[0274] A "recording device" refers to a device that automatically acquires video footage of a user's daily activities based on specified conditions.

[0275] A "compression device" refers to a device that converts acquired video information into a predetermined format so that it can be efficiently stored or transmitted.

[0276] A "transmitting device" refers to a device that transmits compressed video information to a central control unit via a communication network.

[0277] An "analysis device" refers to a device that uses artificial intelligence technology to analyze video information and select specific scenes.

[0278] A "generation device" refers to a device that edits selected video information based on analysis to create visual media.

[0279] A "user interface" refers to a screen or control system that presents generated visual information to the user and allows them to make corrections using an intuitive method.

[0280] A "posting device" refers to a device used to transmit modified visual information media to an electronic platform.

[0281] "Visual information media" refers to video content produced through editing.

[0282] "Artificial intelligence technology" is a general term for computer science methods that understand data, learn patterns, or make predictions.

[0283] The system of this invention begins with a recording device worn by the user. The user carries a small recording device that incorporates a camera and sensors. This recording device, with pre-installed software, monitors the user's activities in real time and automatically acquires video when set conditions are met. For example, the camera is activated when a specific geographical location is reached using a position sensor and clock function. This process utilizes conventional digital camera technology and motion detection algorithms.

[0284] The acquired video information is compressed through a compression device within the recording apparatus. For this compression, standardized video codec technologies such as H.264 and HEVC are used. The compressed data is transferred to the server using wireless communication. For this operation, Wi-Fi and 4G / 5G networks are utilized.

[0285] The server receives the compressed video data transmitted through the network and analyzes the content by an analysis device. For this analysis, artificial intelligence technology is used, particularly deep learning models are used to extract specific scenes and actions within the video. As specific software, frameworks such as TensorFlow and PyTorch can be considered. For example, it is possible to extract only the scenes related to tourist destinations and attractions from the video taken by the user during travel.

[0286] Based on the analysis results, the generation device edits the video. During the editing process, optimal rearrangement of the sequence, deletion of unnecessary parts, and furthermore, mosaic processing to prevent personal identification are executed. Also, for adding subtitles involving character conversion of audio data, natural language processing technology using AI is used. For example, Google Cloud Speech-to-Text, which is speech recognition software, is utilized to convert speech into text.

[0287] The user interface provides the generated visual media to the user and enables preview and manual fine-tuning. This interface is optimized for touch operations and ensures intuitive and simple operability. For example, the user can tap on a scene using a tablet or smartphone and change the position of the sequence by drag & drop.

[0288] Finally, the modified visual media is automatically published to an online platform via the server's posting device. The user can seamlessly share the video to multiple SNS platforms with just a few operations.

[0289] As an example of a prompt, the generative AI model can receive instructions for the system in the form of, "Please describe the process of automatically recording visits to tourist spots during a trip and then automatically editing the video and posting it on social media after returning home."

[0290] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0291] Step 1:

[0292] The user wears a small recording device and goes about their daily activities. This device uses a built-in camera and sensors to detect the user's current location and movements. When pre-set conditions (e.g., a specific location or time) are met, the camera automatically activates and video is captured. The input is the user's movement data and current location, and the output is the raw video data captured when the conditions are met.

[0293] Step 2:

[0294] A compression device built into the recording device compresses the acquired video data. Specifically, it reduces the amount of data using compression technologies such as H.264 and HEVC. The input is raw video data, and the output is a compressed video file. This compression enables efficient data transfer.

[0295] Step 3:

[0296] The compressed video file is transmitted to the server via wireless communication (Wi-Fi or mobile data communication). The terminal uses the network through the transmitting device. The input is compressed video data, and the output is video data securely stored on the server.

[0297] Step 4:

[0298] The server passes the received video data to the analysis device. The analysis device uses a deep learning algorithm with a generative AI model to analyze the video data and automatically extract specific important scenes. The input is compressed video data on the server, and the output is video data with the selected scenes.

[0299] Step 5:

[0300] Based on the analysis results, the server's generation device edits the video. Specific actions include rearranging scenes, applying mosaic effects, and adding subtitles. If audio data is available, natural language processing is used to convert it to text and embed subtitles into the video. The input is the analyzed video data, and the output is the edited visual information medium.

[0301] Step 6:

[0302] The server sends the generated visual information medium to the user's terminal. The user reviews and adjusts the edited content through a dedicated user interface. For example, fine-tuning can be done using sliders or touch controls. The input is the edited visual information medium, and the output is the visual information medium with final adjustments made by the user.

[0303] Step 7:

[0304] Ultimately, the server's posting system publishes the adjusted visual information medium to the designated online platform. Users can select various social networking services and simultaneously distribute the video to multiple platforms with just a few operations. The input is the final adjusted visual information medium, and the output is the published video content.

[0305] (Application Example 1)

[0306] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0307] In modern times, there is an increasing demand for individuals to record their daily activities and then distribute them as high-quality video content. However, conventional methods require a lot of manual work and advanced techniques from recording to editing and distribution, making it difficult for users without specialized skills. Furthermore, it is time-consuming to extract and edit important scenes from the captured video, and there is a problem that content generation and distribution cannot be carried out smoothly.

[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0309] In this invention, the server includes means for automatically capturing video by a recording device based on conditions, means for analyzing video data using AI technology by an analysis device to extract important scenes, and means for distributing user-generated content based on the visual media generated on a smart device. As a result, users can easily record, edit, and distribute their daily activities without the need for specialized knowledge.

[0310] The "recording device" is a device having a function of automatically capturing video based on conditions set by the user.

[0311] The "compression means" is means having a function of converting the captured video data into a predetermined format to enable efficient storage and transmission.

[0312] The "transmission device" is a device for uploading the compressed video data to a central processing unit via a network.

[0313] The "analysis device" is a device having a function of analyzing video data using AI technology to extract important scenes and specific actions.

[0314] The "generation device" is a device equipped with a function of editing the analyzed video data to generate the final visual media.

[0315] A "user interface" is an interface that presents generated visual media to the user and allows them to adjust the visual media.

[0316] A "posting device" is a device for publishing a modified visual medium on an online platform.

[0317] A "smart device" is a device that has the ability to process multimedia content and distribute it over the internet.

[0318] Regarding embodiments for carrying out the invention, this invention is a system that automatically generates and distributes high-quality video content of a user's daily activities. This system involves the coordinated operation of various components, including a recording device, an analysis device, a generation device, a user interface, and a transmission device.

[0319] First, the recording device automatically captures important scenes from daily life based on conditions set by the user. Specifically, the camera detects the user's movements and location information and starts capturing video. Smartphones and wearable devices are often used as recording devices.

[0320] The captured video data is converted to a predetermined format by a compression method, preparing it for efficient transmission. The compressed data is then uploaded by the transmission device to the server's central processing unit via the network.

[0321] The analysis system on the server uses a generative AI model to analyze video data and automatically extract important scenes and actions. It also analyzes audio data as needed and generates text using natural language processing techniques. AI libraries such as TensorFlow and OpenCV are used for this analysis process.

[0322] The analyzed data is passed to a generation device, where it is edited into a visual medium using AI technology. This process automatically performs operations such as mosaic processing and adding text to audio. The generated visual medium is provided to the user through a user interface, which can be adjusted as needed. This interface is designed for intuitive user operation.

[0323] The final visual medium is published on an online platform via a posting device. At this stage, user-generated content is delivered directly from smart devices. For example, in a video themed around a family barbecue, the camera can capture smiles and cooking scenes, and natural conversations and cheers can be transcribed and added to the video. After returning home, this video can be easily reviewed and shared on various platforms.

[0324] An example of a prompt would be: "Generate a scenario that visualizes a family enjoying a barbecue together and displays their particularly enjoyable conversation as text."

[0325] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0326] Step 1:

[0327] The recording device detects the user's movements and location information in real time and determines when the trigger conditions for video capture are met. Based on this, the camera automatically activates and acquires video data. The input is user movement and location data, and the output is captured video data. Specifically, smartphones and wearable devices are used.

[0328] Step 2:

[0329] The compression mechanism on the terminal compresses the captured video data into a predetermined format. The input is the captured video data, and the output is the compressed video data. This process applies a highly efficient compression algorithm to reduce the data size.

[0330] Step 3:

[0331] The terminal's transmitting device uploads compressed video data to the central processing unit via the network. The input is compressed video data, and the output is data stored on the server. This step uses a secure protocol to ensure the reliability of data transfer.

[0332] Step 4:

[0333] The server uses an analysis device to analyze uploaded video data with AI technology and extract important scenes and actions. The input is video data stored on the server, and the output is a list of extracted important scenes. A generative AI model is used for scene recognition and action detection.

[0334] Step 5:

[0335] The server operates the generation device, edits the video based on key scenes, and generates the visual medium. The input is a list of extracted key scenes, and the output is the completed visual medium. Specifically, the order of the scenes is automatically adjusted, and mosaic processing or text addition is performed as needed. Using natural language processing technology, the server analyzes the audio data and overlays the generated text onto the video.

[0336] Step 6:

[0337] The user interface presents the generated visual medium to the user and allows for easy adjustments. The input is the completed visual medium, and the output is the visual medium adjusted by the user. The user can intuitively edit through the interface.

[0338] Step 7:

[0339] The terminal's posting device publishes the customized visual medium to an online platform. The input is the user-customized visual medium, and the output is the published online content. This step seamlessly distributes content across multiple platforms.

[0340] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0341] This invention provides a system that integrates a recording device, an emotion engine, an analysis device, and a generation device, enabling users to automatically generate video content while naturally incorporating their own emotions and actions. Users can easily record their daily activities by wearing a small recording device. The recording device detects the user's movements and location information and automatically captures video based on specific conditions.

[0342] The recorded video data is efficiently compressed using a compression method and then transmitted to a server via the network. The server processes the received video data with an analysis device and uses an emotion engine to recognize emotions from the user's facial expressions and voice. This makes it possible to identify important scenes based on the user's emotions within the video data.

[0343] The generation device takes into account the emotional data recognized by the emotion engine and performs appropriate video editing. Specifically, it customizes the visual medium to match the user's emotions by changing the tone of the video and automatically setting background music according to the user's feelings of joy or sadness. Because this process is fully automated by AI technology, users do not need to perform any technical operations and can use it intuitively.

[0344] The completed visual content is presented to the user through a user interface, allowing them to preview the video and make simple adjustments as needed. Once edited, the posting device automatically publishes the visual content to the online platform, enabling users to share the content with a wide audience with a single click.

[0345] As a concrete example, imagine a user attending a birthday party, and a recording device captures the event. In this case, the emotion engine detects the user's smiles and cheers, and the generator edits the video to a brighter tone based on this, adding cheerful music to the background. Ultimately, a video content optimized for emotion data is generated and can be easily shared.

[0346] The following describes the processing flow.

[0347] Step 1:

[0348] The device, worn by the user, uses sensors to detect the user's movements and surrounding sounds, and automatically begins capturing video and audio when predetermined conditions are met.

[0349] Step 2:

[0350] The device compresses the captured video and audio data into a compressed format for efficient processing and temporarily stores it in its internal storage. The time of capture and location information are also recorded as metadata.

[0351] Step 3:

[0352] The device uploads compressed video data to the server when a network connection becomes available. The appropriate connection method (Wi-Fi or mobile data) is automatically selected.

[0353] Step 4:

[0354] The server uses AI technology to analyze the received video data with an analysis device, and detects important scenes and events within the video.

[0355] Step 5:

[0356] The server utilizes an emotion engine to recognize emotions from the user's facial expressions and voice in the video. For example, emotions such as smiles, surprise, and joy are extracted.

[0357] Step 6:

[0358] The generation device edits the video based on recognized emotions and analyzed data of key scenes. The video's color tone, scene flow, and background music are automatically adjusted to match the user's emotions.

[0359] Step 7:

[0360] The server provides the edited visual media to the user as a preview through the user interface. The user can review the video content and make simple editing requests if necessary.

[0361] Step 8:

[0362] The server generates the final visual medium incorporating user feedback and automatically publishes it to the online platform via the posting device. Users can then easily share the video with their audience.

[0363] (Example 2)

[0364] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0365] Current technology makes it difficult to efficiently record users' daily activities as video and automatically generate customized visual media that responds to their emotions based on that footage. Furthermore, there is a lack of readily available methods for users to intuitively edit videos without requiring technical knowledge and easily share them on online platforms. Additionally, processing that conserves network bandwidth through video data compression and transmission while protecting user privacy is essential.

[0366] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0367] In this invention, the server includes means for an analysis device that uses AI technology to analyze video data and identify important scenes, including the user's emotions; means for a generation device that edits video data based on the emotion data and generates visual media; and means for compression used to enable low-bandwidth communication. This enables efficient recording of daily activities, automatic generation of customized visual media based on emotions, and user-friendly, intuitive video editing and easy content sharing.

[0368] A "recording device" is a device that has the function of automatically acquiring a user's daily actions and events as video.

[0369] "Compression means" refers to a technology that converts captured video data into a predetermined format to reduce the amount of data.

[0370] A "transmission device" is a device that has the function of uploading compressed video data to an information processing device via a communication network.

[0371] An "analysis device" is a device that uses AI technology to analyze video data and identify important scenes, particularly based on emotions.

[0372] A "generation device" is a device that edits video data based on analysis results to generate visual media that matches emotions.

[0373] A "user interface" is a means of presenting a generated visual medium to the user and providing interaction to allow for easy adjustments.

[0374] A "posting device" is a device that has the function of automatically publishing a modified visual medium to an internet platform.

[0375] A "communication network" is an infrastructure for electronically sending and receiving data.

[0376] "AI technology" refers to technologies that use artificial intelligence to perform data analysis and decision-making.

[0377] "Visual media" refers to content that is edited to match the user's emotions and presented visually.

[0378] The system of this invention is a complex system including a recording device, a compression means, a transmission device, an analysis device, a generation device, a user interface, and a posting device. Specific embodiments of each component are shown below.

[0379] Recording device

[0380] Users can wear a small recording device that allows them to naturally record their daily activities. This device uses sensors to detect the user's movements and location, and automatically acquires video under specific conditions (for example, when the user is participating in a particular event).

[0381] Compression means and transmission device

[0382] The terminal receives video data acquired from the recording device, applies a video compression algorithm such as H.264 to compress the data, and then transmits the compressed data to the server via the communication network.

[0383] analysis device

[0384] The server processes the received video data using an analysis device. This analysis utilizes AI technology, particularly deep learning, to analyze the user's facial expressions and voice, and identify their emotions. Based on the analysis results, important scenes within the video are identified.

[0385] generator

[0386] The server automatically edits the video by using the analyzed emotion data in a generation device. This process uses a generation AI model to adjust the video tone and select music according to the user's emotions.

[0387] User interface and posting device

[0388] Users view the edited visual media through the device's user interface. If minor adjustments are needed, they can be made on the interface. Finally, the content is published to an online platform using a posting device and shared with a wide audience.

[0389] As a concrete example, imagine a user attending a friend's birthday party, with the recording device capturing the event. The analysis device recognizes emotions of joy from smiles and tone of voice, and the generation device edits the video into a bright-toned image, adding upbeat background music. The finished content can then be easily shared on social media.

[0390] An example of a prompt message might be, "I want you to automatically generate a video that emphasizes a cheerful atmosphere using footage from a friend's birthday party."

[0391] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0392] Step 1:

[0393] The user acquires video using a small recording device. The recording device uses sensors to detect the user's movements and location, and automatically captures video in specific situations. The input is the user's location information and motion sensor data, and the output is the captured raw video data. As a concrete example, the recording device records the surrounding environment when the user attends a birthday party.

[0394] Step 2:

[0395] The terminal receives raw video data from the recording device. It then applies a video compression algorithm (e.g., H.264) to the received data to reduce its size. The input is raw video data, and the output is compressed video data. Specifically, the terminal compresses the video and prepares it for network transmission.

[0396] Step 3:

[0397] The terminal transmits compressed video data to the server via a communication network. The input is the compressed video data, and the output is the status indicating that the data has been successfully uploaded to the server. Specifically, the terminal uses its internet connection to send data to the server.

[0398] Step 4:

[0399] The server processes the received compressed video data using an analysis device. During this process, AI technology is used to analyze emotions and important scenes within the video. The input is the compressed video data, and the output is the analyzed emotion data and information on important scenes. Specifically, the server analyzes smiles and sounds within the video.

[0400] Step 5:

[0401] The server edits the video based on emotional data analyzed using a generator. Tone adjustments and automatic background music settings are performed. The input is the analyzed emotional data, and the output is the edited visual medium. Specifically, the server adds bright visual effects to match the emotion of joy.

[0402] Step 6:

[0403] The user previews the edited visual medium on the user interface on their device. They can adjust the tone, volume, and other aspects of the visual medium as needed. The input is the edited visual medium, and the output is improvement suggestions based on user feedback. For example, the user can change the brightness using a slider.

[0404] Step 7:

[0405] The device uses a posting device to publish the final visual medium to the online platform. The input is the finalized visual medium, and the output is the published result on the online platform. Specifically, the device clicks the publish button and shares the content.

[0406] (Application Example 2)

[0407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0408] In modern times, recording and viewing daily activities on video has become commonplace. However, conventional technology lacked the means to automatically edit and optimize videos in a way that resonates with the user's emotions. Furthermore, users had to perform the editing themselves, resulting in high technical hurdles and a lack of ease of use. There is a growing need for a system that dynamically adjusts the tone of the video and background sound according to emotions, and that is intuitive for users to use.

[0409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0410] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to extract important scenes based on the user's emotions, and means for a generation device to automatically optimize the tone and background sound of the video using AI technology based on an emotion engine. This allows users to naturally generate and share video content customized to their emotions without requiring any technical operation.

[0411] A "recording device" is a device that has the function of automatically capturing the user's activities and surrounding environment as video based on certain conditions.

[0412] "Compression means" refers to a process that converts captured video data into a predetermined format and efficiently reduces its size.

[0413] A "transmission device" is a device equipped with the function of uploading compressed video data to a central processing unit via a network.

[0414] An "analysis device" is a device that uses AI technology to automatically extract user emotions and important scenes from video data.

[0415] An "emotion engine" refers to artificial intelligence technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0416] A "generation device" is a device that uses AI technology to edit extracted video data and generate visual media optimized according to emotions.

[0417] A "user interface" is an interface that presents generated visual media to the user and allows for previewing and adjustments.

[0418] A "posting device" is a device that has the function of publishing a modified visual medium on an online platform.

[0419] "Eye-tracking information" refers to data that detects the user's eye movements and uses that information to set importance levels.

[0420] A "generative AI model" refers to an artificial intelligence model that creates and edits new video content based on given information.

[0421] A "prompt statement" refers to an input statement used to instruct a generative AI model on how to generate content.

[0422] This invention relates to a system that automatically captures user behavior and emotions and generates optimized video content. The system consists of a recording device, an analysis device, a generation device, a user interface, and a posting device.

[0423] The recording device has the function of automatically capturing video based on conditions and is implemented in the form of a small portable device or smart glasses. For example, when a user is wearing smart glasses, the device continuously measures the surrounding video and audio and captures data as needed.

[0424] The analysis device operates on a server and processes captured video data using AI technology. This device analyzes the user's emotions from facial expressions and voice, and extracts important scenes. It uses Microsoft Azure Face API and Google Cloud Vision API for emotion recognition. For example, in a scene where a user is chatting with a friend in a cafe, the device detects a smile and marks it as a high-priority moment.

[0425] The generation device edits visual media based on emotion analysis results within the server. It utilizes a generation AI model to customize the tone and background sound of extracted scenes based on emotions. The AI ​​model operates according to pre-configured prompts. For example, in a scene where a user is enjoying a walk in a park, it would process based on the prompt: "Edit the video of the scene where the user is enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a brighter tone to the scenery."

[0426] The user interface presents the generated visual media to the user, allowing for previews and simple adjustments. This operation takes place on the user's mobile device, providing an intuitive UI. Users can visually review the generated video and switch scenes or adjust effects as needed.

[0427] Finally, the posting device automatically publishes the edited visual medium to the online platform. This feature allows users to widely share video content with a single click. For example, users can easily post a video compilation of travel memories to social media.

[0428] This system provides users with an environment where they can generate and widely share emotionally-based, customized video content without requiring any technical operation.

[0429] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0430] Step 1:

[0431] When a user puts on smart glasses and begins an activity, the recording device starts capturing video and audio data. The input is real-time video and audio from the user's perspective, and the output is the captured raw data. The recording device selectively records data at specific events or timings based on conditions set by the user.

[0432] Step 2:

[0433] The server receives raw data transmitted from the recording device and converts it into a predetermined format using a compression method. The input is the captured raw data, and the output is compressed data with optimized capacity. This process involves data compression to eliminate unnecessary data and reduce network load.

[0434] Step 3:

[0435] The server passes compressed data to the analysis device, which uses AI technology to perform emotion analysis. The input is compressed video and audio data, and the output is emotion data estimated from the user's facial expressions and voice, as well as the extraction of important scenes that should be emphasized based on those emotions. The analysis device uses the Microsoft Azure Face API to analyze facial expressions and determine the user's emotional state. Specifically, it detects the user's smiles and expressions of surprise and identifies scenes to highlight.

[0436] Step 4:

[0437] The generation device performs video editing using a generation AI model, taking into account the emotion data output from the previous step. The input consists of key scenes and emotion data, and the output is the affected visual medium. In this process, based on prompt statements, the device automatically adjusts the tone of the video and sets background music according to instructions such as, "Edit the video of a scene of someone enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a bright tone to the scenery."

[0438] Step 5:

[0439] The terminal presents the generated visual medium to the user through a user interface and accepts simple adjustments. The input is the generated visual medium, and the output is the medium adjusted based on the user's feedback. The terminal displays the video as a preview and provides an interface that allows the user to change scenes and fine-tune effects as needed.

[0440] Step 6:

[0441] The server automatically publishes the finalized visual media to the online platform via a posting device. The input is the adjusted visual media, and the output is the published video content. In this process, the media is uploaded to a pre-configured platform account, allowing users to easily share the content through social media and other means.

[0442] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0443] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0444] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0445] [Third Embodiment]

[0446] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0447] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0448] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0449] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0450] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0451] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0452] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0453] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0454] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0455] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0456] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0457] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0458] This invention is configured as a system that automatically records and edits a user's daily activities. By wearing a small recording device, the user can easily record various scenes of their daily life. The recording device detects the user's movements and location information and automatically activates the camera based on set conditions. As a result, the user does not need to consciously operate the camera, and can capture video in a natural way.

[0459] The captured video data is effectively compressed within the recording device using compression methods and then transferred to a server via the network. The server receives this data and performs analysis using AI technology with an analysis device. In particular, it is possible to extract specific actions or important scenes from the video and omit unnecessary parts.

[0460] Subsequently, the generation device edits the video based on the analysis results. During the editing process, appropriate cuts and scene order are automatically adjusted, and mosaic processing and text addition to the audio are performed. Here, AI-powered natural language processing is used to automatically generate text content and embed it within the video.

[0461] The generated visual media is provided to the user through a user interface. Here, the user can preview the content and make simple adjustments. This adjustment process is designed with the user experience in mind, and features an intuitive interface.

[0462] The final visual medium is automatically published on online platforms via the posting device. Users can seamlessly share content across multiple social networking platforms with a single operation. This entire process allows users to automatically generate and publish high-quality video blogs without relying on specialized technical skills. A concrete example is a scenario where a recording device captures scenes while a user visits a tourist destination during a trip, and the edited video is immediately published online upon their return home.

[0463] The following describes the processing flow.

[0464] Step 1:

[0465] The device uses sensors to detect the user's movements and location, and when pre-set conditions are met, it automatically activates the camera and starts capturing video.

[0466] Step 2:

[0467] The device converts the captured video into a predetermined compression format (e.g., H.264) and temporarily stores it in its internal storage. At the same time, metadata related to the capture (such as the date and time of capture and location information) is also added.

[0468] Step 3:

[0469] When the terminal detects a stable network connection, it automatically begins sending compressed video data stored on the recording device to the server. At this time, it uploads the data using the most suitable method, either Wi-Fi or mobile data.

[0470] Step 4:

[0471] The server receives the transmitted video data and performs analysis using AI technology through an analysis device. This analysis includes detecting important scenes, voice recognition, and detecting specific objects in the video.

[0472] Step 5:

[0473] The server uses a generation device to edit the video based on the analysis results. During the editing process, the system automatically cuts and rearranges scenes, applies mosaic effects, adds background music, and generates and places subtitles as needed.

[0474] Step 6:

[0475] The server provides the user with a preview of the edited visual media through the user interface. The user can review this preview and send feedback and minor adjustments to the server through the interface.

[0476] Step 7:

[0477] The server makes the necessary changes based on user feedback and then exports the final visual medium as the completed version.

[0478] Step 8:

[0479] The server automatically publishes the final visual version on the designated online platform using a posting device, allowing users to share the content via social media.

[0480] (Example 1)

[0481] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0482] The process of recording daily life and later editing and sharing that footage in high quality is time-consuming and requires specialized skills, making it a significant burden for the average user. Therefore, there is a need for a user-friendly, intuitive, and automated video editing and sharing system.

[0483] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0484] In this invention, the server includes means for a recording device to automatically acquire video based on conditions, means for a compression device to convert the acquired video information into a predetermined format, and means for an analysis device to analyze the video information using artificial intelligence technology and select specific scenes. This makes it possible to automatically extract important scenes from daily video recordings and efficiently edit and transmit them.

[0485] A "recording device" refers to a device that automatically acquires video footage of a user's daily activities based on specified conditions.

[0486] A "compression device" refers to a device that converts acquired video information into a predetermined format so that it can be efficiently stored or transmitted.

[0487] A "transmitting device" refers to a device that transmits compressed video information to a central control unit via a communication network.

[0488] An "analysis device" refers to a device that uses artificial intelligence technology to analyze video information and select specific scenes.

[0489] A "generation device" refers to a device that edits selected video information based on analysis to create visual media.

[0490] A "user interface" refers to a screen or control system that presents generated visual information to the user and allows them to make corrections using an intuitive method.

[0491] A "posting device" refers to a device used to transmit modified visual information media to an electronic platform.

[0492] "Visual information media" refers to video content produced through editing.

[0493] "Artificial intelligence technology" is a general term for computer science methods that understand data, learn patterns, or make predictions.

[0494] The system of this invention begins with a recording device worn by the user. The user carries a small recording device that incorporates a camera and sensors. This recording device, with pre-installed software, monitors the user's activities in real time and automatically acquires video when set conditions are met. For example, the camera is activated when a specific geographical location is reached using a position sensor and clock function. This process utilizes conventional digital camera technology and motion detection algorithms.

[0495] The acquired video information is compressed through a compression device within the recording device. Standardized video codec technologies such as H.264 and HEVC are used for this compression. The compressed data is then transferred to a server using wireless communication. Wi-Fi and 4G / 5G networks are utilized for this operation.

[0496] The server receives compressed video data transmitted over the network and analyzes its contents using an analysis device. This analysis utilizes artificial intelligence technology, particularly deep learning models, to extract specific scenes and actions from the video. Specific software frameworks such as TensorFlow and PyTorch can be used. For example, it's possible to extract only scenes related to tourist destinations or attractions from video footage taken by a user during a trip.

[0497] Based on the analysis results, the generation device edits the video. During the editing process, the sequence is optimally rearranged, unnecessary parts are deleted, and mosaic processing is performed to prevent the identification of individuals. In addition, AI-based natural language processing technology is used to add subtitles that involve converting audio data into text. For example, Google Cloud Speech-to-Text, a speech recognition software, is used to convert the audio into text.

[0498] The user interface provides the user with the generated visual medium, allowing for preview and fine-tuning manually. This interface is optimized for touch operation, ensuring intuitive and easy usability. For example, users can tap scenes on a tablet or smartphone and drag and drop to change the sequence position.

[0499] Finally, the modified visual media is automatically published to online platforms via the server's posting system. Users can seamlessly share the video to multiple social networking platforms with just a few clicks.

[0500] As an example of a prompt, the generative AI model can receive instructions for the system in the form of, "Please describe the process of automatically recording visits to tourist spots during a trip and then automatically editing the video and posting it on social media after returning home."

[0501] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0502] Step 1:

[0503] The user wears a small recording device and goes about their daily activities. This device uses a built-in camera and sensors to detect the user's current location and movements. When pre-set conditions (e.g., a specific location or time) are met, the camera automatically activates and video is captured. The input is the user's movement data and current location, and the output is the raw video data captured when the conditions are met.

[0504] Step 2:

[0505] A compression device built into the recording device compresses the acquired video data. Specifically, it reduces the amount of data using compression technologies such as H.264 and HEVC. The input is raw video data, and the output is a compressed video file. This compression enables efficient data transfer.

[0506] Step 3:

[0507] The compressed video file is transmitted to the server via wireless communication (Wi-Fi or mobile data communication). The terminal uses the network through the transmitting device. The input is compressed video data, and the output is video data securely stored on the server.

[0508] Step 4:

[0509] The server passes the received video data to the analysis device. The analysis device uses a deep learning algorithm with a generative AI model to analyze the video data and automatically extract specific important scenes. The input is compressed video data on the server, and the output is video data with the selected scenes.

[0510] Step 5:

[0511] Based on the analysis results, the server's generation device edits the video. Specific actions include rearranging scenes, applying mosaic effects, and adding subtitles. If audio data is available, natural language processing is used to convert it to text and embed subtitles into the video. The input is the analyzed video data, and the output is the edited visual information medium.

[0512] Step 6:

[0513] The server sends the generated visual information medium to the user's terminal. The user reviews and adjusts the edited content through a dedicated user interface. For example, fine-tuning can be done using sliders or touch controls. The input is the edited visual information medium, and the output is the visual information medium with final adjustments made by the user.

[0514] Step 7:

[0515] Ultimately, the server's posting system publishes the adjusted visual information medium to the designated online platform. Users can select various social networking services and simultaneously distribute the video to multiple platforms with just a few operations. The input is the final adjusted visual information medium, and the output is the published video content.

[0516] (Application Example 1)

[0517] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0518] In modern times, there is a growing demand for individuals to record their daily activities and then distribute them as high-quality video content. However, traditional methods require a lot of manual work and advanced skills from recording to editing and distribution, making it difficult for users without specialized skills. Furthermore, extracting and editing important scenes from the recorded footage is time-consuming, hindering the smooth creation and distribution of content.

[0519] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0520] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to analyze video data using AI technology and extract important scenes, and means for delivering user-generated content based on visual media generated on a smart device. This makes it possible for users to easily record, edit, and distribute their daily activities without requiring specialized knowledge.

[0521] A "recording device" is a device that has the function of automatically capturing video based on conditions set by the user.

[0522] "Compression means" refers to means that have the function of converting captured video data into a predetermined format, making it efficiently stored and transmitted.

[0523] A "transmission device" is a device used to upload compressed video data to a central processing unit via a network.

[0524] An "analysis device" is a device that uses AI technology to analyze video data and extract important scenes or specific actions.

[0525] A "generation device" is a device equipped with the function of editing analyzed video data and generating the final visual medium.

[0526] A "user interface" is an interface that presents generated visual media to the user and allows them to adjust the visual media.

[0527] A "posting device" is a device for publishing a modified visual medium on an online platform.

[0528] A "smart device" is a device that has the ability to process multimedia content and distribute it over the internet.

[0529] Regarding embodiments for carrying out the invention, this invention is a system that automatically generates and distributes high-quality video content of a user's daily activities. This system involves the coordinated operation of various components, including a recording device, an analysis device, a generation device, a user interface, and a transmission device.

[0530] First, the recording device automatically captures important scenes from daily life based on conditions set by the user. Specifically, the camera detects the user's movements and location information and starts capturing video. Smartphones and wearable devices are often used as recording devices.

[0531] The captured video data is converted to a predetermined format by a compression method, preparing it for efficient transmission. The compressed data is then uploaded by the transmission device to the server's central processing unit via the network.

[0532] The analysis system on the server uses a generative AI model to analyze video data and automatically extract important scenes and actions. It also analyzes audio data as needed and generates text using natural language processing techniques. AI libraries such as TensorFlow and OpenCV are used for this analysis process.

[0533] The analyzed data is passed to a generation device, where it is edited into a visual medium using AI technology. This process automatically performs operations such as mosaic processing and adding text to audio. The generated visual medium is provided to the user through a user interface, which can be adjusted as needed. This interface is designed for intuitive user operation.

[0534] The final visual medium is published on an online platform via a posting device. At this stage, user-generated content is delivered directly from smart devices. For example, in a video themed around a family barbecue, the camera can capture smiles and cooking scenes, and natural conversations and cheers can be transcribed and added to the video. After returning home, this video can be easily reviewed and shared on various platforms.

[0535] An example of a prompt would be: "Generate a scenario that visualizes a family enjoying a barbecue together and displays their particularly enjoyable conversation as text."

[0536] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0537] Step 1:

[0538] The recording device detects the user's movements and location information in real time and determines when the trigger conditions for video capture are met. Based on this, the camera automatically activates and acquires video data. The input is user movement and location data, and the output is captured video data. Specifically, smartphones and wearable devices are used.

[0539] Step 2:

[0540] The compression mechanism on the terminal compresses the captured video data into a predetermined format. The input is the captured video data, and the output is the compressed video data. This process applies a highly efficient compression algorithm to reduce the data size.

[0541] Step 3:

[0542] The terminal's transmitting device uploads compressed video data to the central processing unit via the network. The input is compressed video data, and the output is data stored on the server. This step uses a secure protocol to ensure the reliability of data transfer.

[0543] Step 4:

[0544] The server uses an analysis device to analyze uploaded video data with AI technology and extract important scenes and actions. The input is video data stored on the server, and the output is a list of extracted important scenes. A generative AI model is used for scene recognition and action detection.

[0545] Step 5:

[0546] The server operates the generation device, edits the video based on key scenes, and generates the visual medium. The input is a list of extracted key scenes, and the output is the completed visual medium. Specifically, the order of the scenes is automatically adjusted, and mosaic processing or text addition is performed as needed. Using natural language processing technology, the server analyzes the audio data and overlays the generated text onto the video.

[0547] Step 6:

[0548] The user interface presents the generated visual medium to the user and allows for easy adjustments. The input is the completed visual medium, and the output is the visual medium adjusted by the user. The user can intuitively edit through the interface.

[0549] Step 7:

[0550] The terminal's posting device publishes the customized visual medium to an online platform. The input is the user-customized visual medium, and the output is the published online content. This step seamlessly distributes content across multiple platforms.

[0551] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0552] This invention provides a system that integrates a recording device, an emotion engine, an analysis device, and a generation device, enabling users to automatically generate video content while naturally incorporating their own emotions and actions. Users can easily record their daily activities by wearing a small recording device. The recording device detects the user's movements and location information and automatically captures video based on specific conditions.

[0553] The recorded video data is efficiently compressed using a compression method and then transmitted to a server via the network. The server processes the received video data with an analysis device and uses an emotion engine to recognize emotions from the user's facial expressions and voice. This makes it possible to identify important scenes based on the user's emotions within the video data.

[0554] The generation device takes into account the emotional data recognized by the emotion engine and performs appropriate video editing. Specifically, it customizes the visual medium to match the user's emotions by changing the tone of the video and automatically setting background music according to the user's feelings of joy or sadness. Because this process is fully automated by AI technology, users do not need to perform any technical operations and can use it intuitively.

[0555] The completed visual content is presented to the user through a user interface, allowing them to preview the video and make simple adjustments as needed. Once edited, the posting device automatically publishes the visual content to the online platform, enabling users to share the content with a wide audience with a single click.

[0556] As a concrete example, imagine a user attending a birthday party, and a recording device captures the event. In this case, the emotion engine detects the user's smiles and cheers, and the generator edits the video to a brighter tone based on this, adding cheerful music to the background. Ultimately, a video content optimized for emotion data is generated and can be easily shared.

[0557] The following describes the processing flow.

[0558] Step 1:

[0559] The device, worn by the user, uses sensors to detect the user's movements and surrounding sounds, and automatically begins capturing video and audio when predetermined conditions are met.

[0560] Step 2:

[0561] The device compresses the captured video and audio data into a compressed format for efficient processing and temporarily stores it in its internal storage. The time of capture and location information are also recorded as metadata.

[0562] Step 3:

[0563] The device uploads compressed video data to the server when a network connection becomes available. The appropriate connection method (Wi-Fi or mobile data) is automatically selected.

[0564] Step 4:

[0565] The server uses AI technology to analyze the received video data with an analysis device, and detects important scenes and events within the video.

[0566] Step 5:

[0567] The server utilizes an emotion engine to recognize emotions from the user's facial expressions and voice in the video. For example, emotions such as smiles, surprise, and joy are extracted.

[0568] Step 6:

[0569] The generation device edits the video based on recognized emotions and analyzed data of key scenes. The video's color tone, scene flow, and background music are automatically adjusted to match the user's emotions.

[0570] Step 7:

[0571] The server provides the edited visual media to the user as a preview through the user interface. The user can review the video content and make simple editing requests if necessary.

[0572] Step 8:

[0573] The server generates the final visual medium incorporating user feedback and automatically publishes it to the online platform via the posting device. Users can then easily share the video with their audience.

[0574] (Example 2)

[0575] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0576] Current technology makes it difficult to efficiently record users' daily activities as video and automatically generate customized visual media that responds to their emotions based on that footage. Furthermore, there is a lack of readily available methods for users to intuitively edit videos without requiring technical knowledge and easily share them on online platforms. Additionally, processing that conserves network bandwidth through video data compression and transmission while protecting user privacy is essential.

[0577] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0578] In this invention, the server includes means for an analysis device that uses AI technology to analyze video data and identify important scenes, including the user's emotions; means for a generation device that edits video data based on the emotion data and generates visual media; and means for compression used to enable low-bandwidth communication. This enables efficient recording of daily activities, automatic generation of customized visual media based on emotions, and user-friendly, intuitive video editing and easy content sharing.

[0579] A "recording device" is a device that has the function of automatically acquiring a user's daily actions and events as video.

[0580] "Compression means" refers to a technology that converts captured video data into a predetermined format to reduce the amount of data.

[0581] A "transmission device" is a device that has the function of uploading compressed video data to an information processing device via a communication network.

[0582] An "analysis device" is a device that uses AI technology to analyze video data and identify important scenes, particularly based on emotions.

[0583] A "generation device" is a device that edits video data based on analysis results to generate visual media that matches emotions.

[0584] A "user interface" is a means of presenting a generated visual medium to the user and providing interaction to allow for easy adjustments.

[0585] A "posting device" is a device that has the function of automatically publishing a modified visual medium to an internet platform.

[0586] A "communication network" is an infrastructure for electronically sending and receiving data.

[0587] "AI technology" refers to technologies that use artificial intelligence to perform data analysis and decision-making.

[0588] "Visual media" refers to content that is edited to match the user's emotions and presented visually.

[0589] The system of this invention is a complex system including a recording device, a compression means, a transmission device, an analysis device, a generation device, a user interface, and a posting device. Specific embodiments of each component are shown below.

[0590] Recording device

[0591] Users can wear a small recording device that allows them to naturally record their daily activities. This device uses sensors to detect the user's movements and location, and automatically acquires video under specific conditions (for example, when the user is participating in a particular event).

[0592] Compression means and transmission device

[0593] The terminal receives video data acquired from the recording device, applies a video compression algorithm such as H.264 to compress the data, and then transmits the compressed data to the server via the communication network.

[0594] analysis device

[0595] The server processes the received video data using an analysis device. This analysis utilizes AI technology, particularly deep learning, to analyze the user's facial expressions and voice, and identify their emotions. Based on the analysis results, important scenes within the video are identified.

[0596] generator

[0597] The server automatically edits the video by using the analyzed emotion data in a generation device. This process uses a generation AI model to adjust the video tone and select music according to the user's emotions.

[0598] User interface and posting device

[0599] Users view the edited visual media through the device's user interface. If minor adjustments are needed, they can be made on the interface. Finally, the content is published to an online platform using a posting device and shared with a wide audience.

[0600] As a concrete example, imagine a user attending a friend's birthday party, with the recording device capturing the event. The analysis device recognizes emotions of joy from smiles and tone of voice, and the generation device edits the video into a bright-toned image, adding upbeat background music. The finished content can then be easily shared on social media.

[0601] An example of a prompt message might be, "I want you to automatically generate a video that emphasizes a cheerful atmosphere using footage from a friend's birthday party."

[0602] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0603] Step 1:

[0604] The user acquires video using a small recording device. The recording device uses sensors to detect the user's movements and location, and automatically captures video in specific situations. The input is the user's location information and motion sensor data, and the output is the captured raw video data. As a concrete example, the recording device records the surrounding environment when the user attends a birthday party.

[0605] Step 2:

[0606] The terminal receives raw video data from the recording device. It then applies a video compression algorithm (e.g., H.264) to the received data to reduce its size. The input is raw video data, and the output is compressed video data. Specifically, the terminal compresses the video and prepares it for network transmission.

[0607] Step 3:

[0608] The terminal transmits compressed video data to the server via a communication network. The input is the compressed video data, and the output is the status indicating that the data has been successfully uploaded to the server. Specifically, the terminal uses its internet connection to send data to the server.

[0609] Step 4:

[0610] The server processes the received compressed video data using an analysis device. During this process, AI technology is used to analyze emotions and important scenes within the video. The input is the compressed video data, and the output is the analyzed emotion data and information on important scenes. Specifically, the server analyzes smiles and sounds within the video.

[0611] Step 5:

[0612] The server edits the video based on emotional data analyzed using a generator. Tone adjustments and automatic background music settings are performed. The input is the analyzed emotional data, and the output is the edited visual medium. Specifically, the server adds bright visual effects to match the emotion of joy.

[0613] Step 6:

[0614] The user previews the edited visual medium on the user interface on their device. They can adjust the tone, volume, and other aspects of the visual medium as needed. The input is the edited visual medium, and the output is improvement suggestions based on user feedback. For example, the user can change the brightness using a slider.

[0615] Step 7:

[0616] The device uses a posting device to publish the final visual medium to the online platform. The input is the finalized visual medium, and the output is the published result on the online platform. Specifically, the device clicks the publish button and shares the content.

[0617] (Application Example 2)

[0618] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0619] In modern times, recording and viewing daily activities on video has become commonplace. However, conventional technology lacked the means to automatically edit and optimize videos in a way that resonates with the user's emotions. Furthermore, users had to perform the editing themselves, resulting in high technical hurdles and a lack of ease of use. There is a growing need for a system that dynamically adjusts the tone of the video and background sound according to emotions, and that is intuitive for users to use.

[0620] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0621] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to extract important scenes based on the user's emotions, and means for a generation device to automatically optimize the tone and background sound of the video using AI technology based on an emotion engine. This allows users to naturally generate and share video content customized to their emotions without requiring any technical operation.

[0622] A "recording device" is a device that has the function of automatically capturing the user's activities and surrounding environment as video based on certain conditions.

[0623] "Compression means" refers to a process that converts captured video data into a predetermined format and efficiently reduces its size.

[0624] A "transmission device" is a device equipped with the function of uploading compressed video data to a central processing unit via a network.

[0625] An "analysis device" is a device that uses AI technology to automatically extract user emotions and important scenes from video data.

[0626] An "emotion engine" refers to artificial intelligence technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0627] A "generation device" is a device that uses AI technology to edit extracted video data and generate visual media optimized according to emotions.

[0628] A "user interface" is an interface that presents generated visual media to the user and allows for previewing and adjustments.

[0629] A "posting device" is a device that has the function of publishing a modified visual medium on an online platform.

[0630] "Eye-tracking information" refers to data that detects the user's eye movements and uses that information to set importance levels.

[0631] A "generative AI model" refers to an artificial intelligence model that creates and edits new video content based on given information.

[0632] A "prompt statement" refers to an input statement used to instruct a generative AI model on how to generate content.

[0633] This invention relates to a system that automatically captures user behavior and emotions and generates optimized video content. The system consists of a recording device, an analysis device, a generation device, a user interface, and a posting device.

[0634] The recording device has the function of automatically capturing video based on conditions and is implemented in the form of a small portable device or smart glasses. For example, when a user is wearing smart glasses, the device continuously measures the surrounding video and audio and captures data as needed.

[0635] The analysis device operates on a server and processes captured video data using AI technology. This device analyzes the user's emotions from facial expressions and voice, and extracts important scenes. It uses Microsoft Azure Face API and Google Cloud Vision API for emotion recognition. For example, in a scene where a user is chatting with a friend in a cafe, the device detects a smile and marks it as a high-priority moment.

[0636] The generation device edits visual media based on emotion analysis results within the server. It utilizes a generation AI model to customize the tone and background sound of extracted scenes based on emotions. The AI ​​model operates according to pre-configured prompts. For example, in a scene where a user is enjoying a walk in a park, it would process based on the prompt: "Edit the video of the scene where the user is enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a brighter tone to the scenery."

[0637] The user interface presents the generated visual media to the user, allowing for previews and simple adjustments. This operation takes place on the user's mobile device, providing an intuitive UI. Users can visually review the generated video and switch scenes or adjust effects as needed.

[0638] Finally, the posting device automatically publishes the edited visual medium to the online platform. This feature allows users to widely share video content with a single click. For example, users can easily post a video compilation of travel memories to social media.

[0639] This system provides users with an environment where they can generate and widely share emotionally-based, customized video content without requiring any technical operation.

[0640] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0641] Step 1:

[0642] When a user puts on smart glasses and begins an activity, the recording device starts capturing video and audio data. The input is real-time video and audio from the user's perspective, and the output is the captured raw data. The recording device selectively records data at specific events or timings based on conditions set by the user.

[0643] Step 2:

[0644] The server receives raw data transmitted from the recording device and converts it into a predetermined format using a compression method. The input is the captured raw data, and the output is compressed data with optimized capacity. This process involves data compression to eliminate unnecessary data and reduce network load.

[0645] Step 3:

[0646] The server passes compressed data to the analysis device, which uses AI technology to perform emotion analysis. The input is compressed video and audio data, and the output is emotion data estimated from the user's facial expressions and voice, as well as the extraction of important scenes that should be emphasized based on those emotions. The analysis device uses the Microsoft Azure Face API to analyze facial expressions and determine the user's emotional state. Specifically, it detects the user's smiles and expressions of surprise and identifies scenes to highlight.

[0647] Step 4:

[0648] The generation device performs video editing using a generation AI model, taking into account the emotion data output from the previous step. The input consists of key scenes and emotion data, and the output is the affected visual medium. In this process, based on prompt statements, the device automatically adjusts the tone of the video and sets background music according to instructions such as, "Edit the video of a scene of someone enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a bright tone to the scenery."

[0649] Step 5:

[0650] The terminal presents the generated visual medium to the user through a user interface and accepts simple adjustments. The input is the generated visual medium, and the output is the medium adjusted based on the user's feedback. The terminal displays the video as a preview and provides an interface that allows the user to change scenes and fine-tune effects as needed.

[0651] Step 6:

[0652] The server automatically publishes the finalized visual media to the online platform via a posting device. The input is the adjusted visual media, and the output is the published video content. In this process, the media is uploaded to a pre-configured platform account, allowing users to easily share the content through social media and other means.

[0653] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0654] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0655] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0656] [Fourth Embodiment]

[0657] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0658] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0659] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0660] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0661] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0662] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0663] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0664] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0665] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0666] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0667] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0668] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0669] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0670] This invention is configured as a system that automatically records and edits a user's daily activities. By wearing a small recording device, the user can easily record various scenes of their daily life. The recording device detects the user's movements and location information and automatically activates the camera based on set conditions. As a result, the user does not need to consciously operate the camera, and can capture video in a natural way.

[0671] The captured video data is effectively compressed within the recording device using compression methods and then transferred to a server via the network. The server receives this data and performs analysis using AI technology with an analysis device. In particular, it is possible to extract specific actions or important scenes from the video and omit unnecessary parts.

[0672] Subsequently, the generation device edits the video based on the analysis results. During the editing process, appropriate cuts and scene order are automatically adjusted, and mosaic processing and text addition to the audio are performed. Here, AI-powered natural language processing is used to automatically generate text content and embed it within the video.

[0673] The generated visual media is provided to the user through a user interface. Here, the user can preview the content and make simple adjustments. This adjustment process is designed with the user experience in mind, and features an intuitive interface.

[0674] The final visual medium is automatically published on online platforms via the posting device. Users can seamlessly share content across multiple social networking platforms with a single operation. This entire process allows users to automatically generate and publish high-quality video blogs without relying on specialized technical skills. A concrete example is a scenario where a recording device captures scenes while a user visits a tourist destination during a trip, and the edited video is immediately published online upon their return home.

[0675] The following describes the processing flow.

[0676] Step 1:

[0677] The device uses sensors to detect the user's movements and location, and when pre-set conditions are met, it automatically activates the camera and starts capturing video.

[0678] Step 2:

[0679] The device converts the captured video into a predetermined compression format (e.g., H.264) and temporarily stores it in its internal storage. At the same time, metadata related to the capture (such as the date and time of capture and location information) is also added.

[0680] Step 3:

[0681] When the terminal detects a stable network connection, it automatically begins sending compressed video data stored on the recording device to the server. At this time, it uploads the data using the most suitable method, either Wi-Fi or mobile data.

[0682] Step 4:

[0683] The server receives the transmitted video data and performs analysis using AI technology through an analysis device. This analysis includes detecting important scenes, voice recognition, and detecting specific objects in the video.

[0684] Step 5:

[0685] The server uses a generation device to edit the video based on the analysis results. During the editing process, the system automatically cuts and rearranges scenes, applies mosaic effects, adds background music, and generates and places subtitles as needed.

[0686] Step 6:

[0687] The server provides the user with a preview of the edited visual media through the user interface. The user can review this preview and send feedback and minor adjustments to the server through the interface.

[0688] Step 7:

[0689] The server makes the necessary changes based on user feedback and then exports the final visual medium as the completed version.

[0690] Step 8:

[0691] The server automatically publishes the final visual version on the designated online platform using a posting device, allowing users to share the content via social media.

[0692] (Example 1)

[0693] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0694] The process of recording daily life and later editing and sharing that footage in high quality is time-consuming and requires specialized skills, making it a significant burden for the average user. Therefore, there is a need for a user-friendly, intuitive, and automated video editing and sharing system.

[0695] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0696] In this invention, the server includes means for a recording device to automatically acquire video based on conditions, means for a compression device to convert the acquired video information into a predetermined format, and means for an analysis device to analyze the video information using artificial intelligence technology and select specific scenes. This makes it possible to automatically extract important scenes from daily video recordings and efficiently edit and transmit them.

[0697] A "recording device" refers to a device that automatically acquires video footage of a user's daily activities based on specified conditions.

[0698] A "compression device" refers to a device that converts acquired video information into a predetermined format so that it can be efficiently stored or transmitted.

[0699] A "transmitting device" refers to a device that transmits compressed video information to a central control unit via a communication network.

[0700] An "analysis device" refers to a device that uses artificial intelligence technology to analyze video information and select specific scenes.

[0701] A "generation device" refers to a device that edits selected video information based on analysis to create visual media.

[0702] A "user interface" refers to a screen or control system that presents generated visual information to the user and allows them to make corrections using an intuitive method.

[0703] A "posting device" refers to a device used to transmit modified visual information media to an electronic platform.

[0704] "Visual information media" refers to video content produced through editing.

[0705] "Artificial intelligence technology" is a general term for computer science methods that understand data, learn patterns, or make predictions.

[0706] The system of this invention begins with a recording device worn by the user. The user carries a small recording device that incorporates a camera and sensors. This recording device, with pre-installed software, monitors the user's activities in real time and automatically acquires video when set conditions are met. For example, the camera is activated when a specific geographical location is reached using a position sensor and clock function. This process utilizes conventional digital camera technology and motion detection algorithms.

[0707] The acquired video information is compressed through a compression device within the recording device. Standardized video codec technologies such as H.264 and HEVC are used for this compression. The compressed data is then transferred to a server using wireless communication. Wi-Fi and 4G / 5G networks are utilized for this operation.

[0708] The server receives compressed video data transmitted over the network and analyzes its contents using an analysis device. This analysis utilizes artificial intelligence technology, particularly deep learning models, to extract specific scenes and actions from the video. Specific software frameworks such as TensorFlow and PyTorch can be used. For example, it's possible to extract only scenes related to tourist destinations or attractions from video footage taken by a user during a trip.

[0709] Based on the analysis results, the generation device edits the video. During the editing process, the sequence is optimally rearranged, unnecessary parts are deleted, and mosaic processing is performed to prevent the identification of individuals. In addition, AI-based natural language processing technology is used to add subtitles that involve converting audio data into text. For example, Google Cloud Speech-to-Text, a speech recognition software, is used to convert the audio into text.

[0710] The user interface provides the user with the generated visual medium, allowing for preview and fine-tuning manually. This interface is optimized for touch operation, ensuring intuitive and easy usability. For example, users can tap scenes on a tablet or smartphone and drag and drop to change the sequence position.

[0711] Finally, the modified visual media is automatically published to online platforms via the server's posting system. Users can seamlessly share the video to multiple social networking platforms with just a few clicks.

[0712] As an example of a prompt, the generative AI model can receive instructions for the system in the form of, "Please describe the process of automatically recording visits to tourist spots during a trip and then automatically editing the video and posting it on social media after returning home."

[0713] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0714] Step 1:

[0715] The user wears a small recording device and goes about their daily activities. This device uses a built-in camera and sensors to detect the user's current location and movements. When pre-set conditions (e.g., a specific location or time) are met, the camera automatically activates and video is captured. The input is the user's movement data and current location, and the output is the raw video data captured when the conditions are met.

[0716] Step 2:

[0717] A compression device built into the recording device compresses the acquired video data. Specifically, it reduces the amount of data using compression technologies such as H.264 and HEVC. The input is raw video data, and the output is a compressed video file. This compression enables efficient data transfer.

[0718] Step 3:

[0719] The compressed video file is transmitted to the server via wireless communication (Wi-Fi or mobile data communication). The terminal uses the network through the transmitting device. The input is compressed video data, and the output is video data securely stored on the server.

[0720] Step 4:

[0721] The server passes the received video data to the analysis device. The analysis device uses a deep learning algorithm with a generative AI model to analyze the video data and automatically extract specific important scenes. The input is compressed video data on the server, and the output is video data with the selected scenes.

[0722] Step 5:

[0723] Based on the analysis results, the server's generation device edits the video. Specific actions include rearranging scenes, applying mosaic effects, and adding subtitles. If audio data is available, natural language processing is used to convert it to text and embed subtitles into the video. The input is the analyzed video data, and the output is the edited visual information medium.

[0724] Step 6:

[0725] The server sends the generated visual information medium to the user's terminal. The user reviews and adjusts the edited content through a dedicated user interface. For example, fine-tuning can be done using sliders or touch controls. The input is the edited visual information medium, and the output is the visual information medium with final adjustments made by the user.

[0726] Step 7:

[0727] Ultimately, the server's posting system publishes the adjusted visual information medium to the designated online platform. Users can select various social networking services and simultaneously distribute the video to multiple platforms with just a few operations. The input is the final adjusted visual information medium, and the output is the published video content.

[0728] (Application Example 1)

[0729] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0730] In modern times, there is a growing demand for individuals to record their daily activities and then distribute them as high-quality video content. However, traditional methods require a lot of manual work and advanced skills from recording to editing and distribution, making it difficult for users without specialized skills. Furthermore, extracting and editing important scenes from the recorded footage is time-consuming, hindering the smooth creation and distribution of content.

[0731] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0732] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to analyze video data using AI technology and extract important scenes, and means for delivering user-generated content based on visual media generated on a smart device. This makes it possible for users to easily record, edit, and distribute their daily activities without requiring specialized knowledge.

[0733] A "recording device" is a device that has the function of automatically capturing video based on conditions set by the user.

[0734] "Compression means" refers to means that have the function of converting captured video data into a predetermined format, making it efficiently stored and transmitted.

[0735] A "transmission device" is a device used to upload compressed video data to a central processing unit via a network.

[0736] An "analysis device" is a device that uses AI technology to analyze video data and extract important scenes or specific actions.

[0737] A "generation device" is a device equipped with the function of editing analyzed video data and generating the final visual medium.

[0738] A "user interface" is an interface that presents generated visual media to the user and allows them to adjust the visual media.

[0739] A "posting device" is a device for publishing a modified visual medium on an online platform.

[0740] A "smart device" is a device that has the ability to process multimedia content and distribute it over the internet.

[0741] Regarding embodiments for carrying out the invention, this invention is a system that automatically generates and distributes high-quality video content of a user's daily activities. This system involves the coordinated operation of various components, including a recording device, an analysis device, a generation device, a user interface, and a transmission device.

[0742] First, the recording device automatically captures important scenes from daily life based on conditions set by the user. Specifically, the camera detects the user's movements and location information and starts capturing video. Smartphones and wearable devices are often used as recording devices.

[0743] The captured video data is converted to a predetermined format by a compression method, preparing it for efficient transmission. The compressed data is then uploaded by the transmission device to the server's central processing unit via the network.

[0744] The analysis system on the server uses a generative AI model to analyze video data and automatically extract important scenes and actions. It also analyzes audio data as needed and generates text using natural language processing techniques. AI libraries such as TensorFlow and OpenCV are used for this analysis process.

[0745] The analyzed data is passed to a generation device, where it is edited into a visual medium using AI technology. This process automatically performs operations such as mosaic processing and adding text to audio. The generated visual medium is provided to the user through a user interface, which can be adjusted as needed. This interface is designed for intuitive user operation.

[0746] The final visual medium is published on an online platform via a posting device. At this stage, user-generated content is delivered directly from smart devices. For example, in a video themed around a family barbecue, the camera can capture smiles and cooking scenes, and natural conversations and cheers can be transcribed and added to the video. After returning home, this video can be easily reviewed and shared on various platforms.

[0747] An example of a prompt would be: "Generate a scenario that visualizes a family enjoying a barbecue together and displays their particularly enjoyable conversation as text."

[0748] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0749] Step 1:

[0750] The recording device detects the user's movements and location information in real time and determines when the trigger conditions for video capture are met. Based on this, the camera automatically activates and acquires video data. The input is user movement and location data, and the output is captured video data. Specifically, smartphones and wearable devices are used.

[0751] Step 2:

[0752] The compression mechanism on the terminal compresses the captured video data into a predetermined format. The input is the captured video data, and the output is the compressed video data. This process applies a highly efficient compression algorithm to reduce the data size.

[0753] Step 3:

[0754] The terminal's transmitting device uploads compressed video data to the central processing unit via the network. The input is compressed video data, and the output is data stored on the server. This step uses a secure protocol to ensure the reliability of data transfer.

[0755] Step 4:

[0756] The server uses an analysis device to analyze uploaded video data with AI technology and extract important scenes and actions. The input is video data stored on the server, and the output is a list of extracted important scenes. A generative AI model is used for scene recognition and action detection.

[0757] Step 5:

[0758] The server operates the generation device, edits the video based on key scenes, and generates the visual medium. The input is a list of extracted key scenes, and the output is the completed visual medium. Specifically, the order of the scenes is automatically adjusted, and mosaic processing or text addition is performed as needed. Using natural language processing technology, the server analyzes the audio data and overlays the generated text onto the video.

[0759] Step 6:

[0760] The user interface presents the generated visual medium to the user and allows for easy adjustments. The input is the completed visual medium, and the output is the visual medium adjusted by the user. The user can intuitively edit through the interface.

[0761] Step 7:

[0762] The terminal's posting device publishes the customized visual medium to an online platform. The input is the user-customized visual medium, and the output is the published online content. This step seamlessly distributes content across multiple platforms.

[0763] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0764] This invention provides a system that integrates a recording device, an emotion engine, an analysis device, and a generation device, enabling users to automatically generate video content while naturally incorporating their own emotions and actions. Users can easily record their daily activities by wearing a small recording device. The recording device detects the user's movements and location information and automatically captures video based on specific conditions.

[0765] The recorded video data is efficiently compressed using a compression method and then transmitted to a server via the network. The server processes the received video data with an analysis device and uses an emotion engine to recognize emotions from the user's facial expressions and voice. This makes it possible to identify important scenes based on the user's emotions within the video data.

[0766] The generation device takes into account the emotional data recognized by the emotion engine and performs appropriate video editing. Specifically, it customizes the visual medium to match the user's emotions by changing the tone of the video and automatically setting background music according to the user's feelings of joy or sadness. Because this process is fully automated by AI technology, users do not need to perform any technical operations and can use it intuitively.

[0767] The completed visual content is presented to the user through a user interface, allowing them to preview the video and make simple adjustments as needed. Once edited, the posting device automatically publishes the visual content to the online platform, enabling users to share the content with a wide audience with a single click.

[0768] As a concrete example, imagine a user attending a birthday party, and a recording device captures the event. In this case, the emotion engine detects the user's smiles and cheers, and the generator edits the video to a brighter tone based on this, adding cheerful music to the background. Ultimately, a video content optimized for emotion data is generated and can be easily shared.

[0769] The following describes the processing flow.

[0770] Step 1:

[0771] The device, worn by the user, uses sensors to detect the user's movements and surrounding sounds, and automatically begins capturing video and audio when predetermined conditions are met.

[0772] Step 2:

[0773] The device compresses the captured video and audio data into a compressed format for efficient processing and temporarily stores it in its internal storage. The time of capture and location information are also recorded as metadata.

[0774] Step 3:

[0775] The device uploads compressed video data to the server when a network connection becomes available. The appropriate connection method (Wi-Fi or mobile data) is automatically selected.

[0776] Step 4:

[0777] The server uses AI technology to analyze the received video data with an analysis device, and detects important scenes and events within the video.

[0778] Step 5:

[0779] The server utilizes an emotion engine to recognize emotions from the user's facial expressions and voice in the video. For example, emotions such as smiles, surprise, and joy are extracted.

[0780] Step 6:

[0781] The generation device edits the video based on recognized emotions and analyzed data of key scenes. The video's color tone, scene flow, and background music are automatically adjusted to match the user's emotions.

[0782] Step 7:

[0783] The server provides the edited visual media to the user as a preview through the user interface. The user can review the video content and make simple editing requests if necessary.

[0784] Step 8:

[0785] The server generates the final visual medium incorporating user feedback and automatically publishes it to the online platform via the posting device. Users can then easily share the video with their audience.

[0786] (Example 2)

[0787] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0788] Current technology makes it difficult to efficiently record users' daily activities as video and automatically generate customized visual media that responds to their emotions based on that footage. Furthermore, there is a lack of readily available methods for users to intuitively edit videos without requiring technical knowledge and easily share them on online platforms. Additionally, processing that conserves network bandwidth through video data compression and transmission while protecting user privacy is essential.

[0789] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0790] In this invention, the server includes means for an analysis device that uses AI technology to analyze video data and identify important scenes, including the user's emotions; means for a generation device that edits video data based on the emotion data and generates visual media; and means for compression used to enable low-bandwidth communication. This enables efficient recording of daily activities, automatic generation of customized visual media based on emotions, and user-friendly, intuitive video editing and easy content sharing.

[0791] A "recording device" is a device that has the function of automatically acquiring a user's daily actions and events as video.

[0792] "Compression means" refers to a technology that converts captured video data into a predetermined format to reduce the amount of data.

[0793] A "transmission device" is a device that has the function of uploading compressed video data to an information processing device via a communication network.

[0794] An "analysis device" is a device that uses AI technology to analyze video data and identify important scenes, particularly based on emotions.

[0795] A "generation device" is a device that edits video data based on analysis results to generate visual media that matches emotions.

[0796] A "user interface" is a means of presenting a generated visual medium to the user and providing interaction to allow for easy adjustments.

[0797] A "posting device" is a device that has the function of automatically publishing a modified visual medium to an internet platform.

[0798] A "communication network" is an infrastructure for electronically sending and receiving data.

[0799] "AI technology" refers to technologies that use artificial intelligence to perform data analysis and decision-making.

[0800] "Visual media" refers to content that is edited to match the user's emotions and presented visually.

[0801] The system of this invention is a complex system including a recording device, a compression means, a transmission device, an analysis device, a generation device, a user interface, and a posting device. Specific embodiments of each component are shown below.

[0802] Recording device

[0803] Users can wear a small recording device that allows them to naturally record their daily activities. This device uses sensors to detect the user's movements and location, and automatically acquires video under specific conditions (for example, when the user is participating in a particular event).

[0804] Compression means and transmission device

[0805] The terminal receives video data acquired from the recording device, applies a video compression algorithm such as H.264 to compress the data, and then transmits the compressed data to the server via the communication network.

[0806] analysis device

[0807] The server processes the received video data using an analysis device. This analysis utilizes AI technology, particularly deep learning, to analyze the user's facial expressions and voice, and identify their emotions. Based on the analysis results, important scenes within the video are identified.

[0808] generator

[0809] The server automatically edits the video by using the analyzed emotion data in a generation device. This process uses a generation AI model to adjust the video tone and select music according to the user's emotions.

[0810] User interface and posting device

[0811] Users view the edited visual media through the device's user interface. If minor adjustments are needed, they can be made on the interface. Finally, the content is published to an online platform using a posting device and shared with a wide audience.

[0812] As a concrete example, imagine a user attending a friend's birthday party, with the recording device capturing the event. The analysis device recognizes emotions of joy from smiles and tone of voice, and the generation device edits the video into a bright-toned image, adding upbeat background music. The finished content can then be easily shared on social media.

[0813] An example of a prompt message might be, "I want you to automatically generate a video that emphasizes a cheerful atmosphere using footage from a friend's birthday party."

[0814] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0815] Step 1:

[0816] The user acquires video using a small recording device. The recording device uses sensors to detect the user's movements and location, and automatically captures video in specific situations. The input is the user's location information and motion sensor data, and the output is the captured raw video data. As a concrete example, the recording device records the surrounding environment when the user attends a birthday party.

[0817] Step 2:

[0818] The terminal receives raw video data from the recording device. It then applies a video compression algorithm (e.g., H.264) to the received data to reduce its size. The input is raw video data, and the output is compressed video data. Specifically, the terminal compresses the video and prepares it for network transmission.

[0819] Step 3:

[0820] The terminal transmits compressed video data to the server via a communication network. The input is the compressed video data, and the output is the status indicating that the data has been successfully uploaded to the server. Specifically, the terminal uses its internet connection to send data to the server.

[0821] Step 4:

[0822] The server processes the received compressed video data using an analysis device. During this process, AI technology is used to analyze emotions and important scenes within the video. The input is the compressed video data, and the output is the analyzed emotion data and information on important scenes. Specifically, the server analyzes smiles and sounds within the video.

[0823] Step 5:

[0824] The server edits the video based on emotional data analyzed using a generator. Tone adjustments and automatic background music settings are performed. The input is the analyzed emotional data, and the output is the edited visual medium. Specifically, the server adds bright visual effects to match the emotion of joy.

[0825] Step 6:

[0826] The user previews the edited visual medium on the user interface on their device. They can adjust the tone, volume, and other aspects of the visual medium as needed. The input is the edited visual medium, and the output is improvement suggestions based on user feedback. For example, the user can change the brightness using a slider.

[0827] Step 7:

[0828] The device uses a posting device to publish the final visual medium to the online platform. The input is the finalized visual medium, and the output is the published result on the online platform. Specifically, the device clicks the publish button and shares the content.

[0829] (Application Example 2)

[0830] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0831] In modern times, recording and viewing daily activities on video has become commonplace. However, conventional technology lacked the means to automatically edit and optimize videos in a way that resonates with the user's emotions. Furthermore, users had to perform the editing themselves, resulting in high technical hurdles and a lack of ease of use. There is a growing need for a system that dynamically adjusts the tone of the video and background sound according to emotions, and that is intuitive for users to use.

[0832] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0833] In this invention, the server includes means for a recording device to automatically capture video based on conditions, means for an analysis device to extract important scenes based on the user's emotions, and means for a generation device to automatically optimize the tone and background sound of the video using AI technology based on an emotion engine. This allows users to naturally generate and share video content customized to their emotions without requiring any technical operation.

[0834] A "recording device" is a device that has the function of automatically capturing the user's activities and surrounding environment as video based on certain conditions.

[0835] "Compression means" refers to a process that converts captured video data into a predetermined format and efficiently reduces its size.

[0836] A "transmission device" is a device equipped with the function of uploading compressed video data to a central processing unit via a network.

[0837] An "analysis device" is a device that uses AI technology to automatically extract user emotions and important scenes from video data.

[0838] An "emotion engine" refers to artificial intelligence technology that analyzes a user's facial expressions and voice to identify their emotional state.

[0839] A "generation device" is a device that uses AI technology to edit extracted video data and generate visual media optimized according to emotions.

[0840] A "user interface" is an interface that presents generated visual media to the user and allows for previewing and adjustments.

[0841] A "posting device" is a device that has the function of publishing a modified visual medium on an online platform.

[0842] "Eye-tracking information" refers to data that detects the user's eye movements and uses that information to set importance levels.

[0843] A "generative AI model" refers to an artificial intelligence model that creates and edits new video content based on given information.

[0844] A "prompt statement" refers to an input statement used to instruct a generative AI model on how to generate content.

[0845] This invention relates to a system that automatically captures user behavior and emotions and generates optimized video content. The system consists of a recording device, an analysis device, a generation device, a user interface, and a posting device.

[0846] The recording device has the function of automatically capturing video based on conditions and is implemented in the form of a small portable device or smart glasses. For example, when a user is wearing smart glasses, the device continuously measures the surrounding video and audio and captures data as needed.

[0847] The analysis device operates on a server and processes captured video data using AI technology. This device analyzes the user's emotions from facial expressions and voice, and extracts important scenes. It uses Microsoft Azure Face API and Google Cloud Vision API for emotion recognition. For example, in a scene where a user is chatting with a friend in a cafe, the device detects a smile and marks it as a high-priority moment.

[0848] The generation device edits visual media based on emotion analysis results within the server. It utilizes a generation AI model to customize the tone and background sound of extracted scenes based on emotions. The AI ​​model operates according to pre-configured prompts. For example, in a scene where a user is enjoying a walk in a park, it would process based on the prompt: "Edit the video of the scene where the user is enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a brighter tone to the scenery."

[0849] The user interface presents the generated visual media to the user, allowing for previews and simple adjustments. This operation takes place on the user's mobile device, providing an intuitive UI. Users can visually review the generated video and switch scenes or adjust effects as needed.

[0850] Finally, the posting device automatically publishes the edited visual medium to the online platform. This feature allows users to widely share video content with a single click. For example, users can easily post a video compilation of travel memories to social media.

[0851] This system provides users with an environment where they can generate and widely share emotionally-based, customized video content without requiring any technical operation.

[0852] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0853] Step 1:

[0854] When a user puts on smart glasses and begins an activity, the recording device starts capturing video and audio data. The input is real-time video and audio from the user's perspective, and the output is the captured raw data. The recording device selectively records data at specific events or timings based on conditions set by the user.

[0855] Step 2:

[0856] The server receives raw data transmitted from the recording device and converts it into a predetermined format using a compression method. The input is the captured raw data, and the output is compressed data with optimized capacity. This process involves data compression to eliminate unnecessary data and reduce network load.

[0857] Step 3:

[0858] The server passes compressed data to the analysis device, which uses AI technology to perform emotion analysis. The input is compressed video and audio data, and the output is emotion data estimated from the user's facial expressions and voice, as well as the extraction of important scenes that should be emphasized based on those emotions. The analysis device uses the Microsoft Azure Face API to analyze facial expressions and determine the user's emotional state. Specifically, it detects the user's smiles and expressions of surprise and identifies scenes to highlight.

[0859] Step 4:

[0860] The generation device performs video editing using a generation AI model, taking into account the emotion data output from the previous step. The input consists of key scenes and emotion data, and the output is the affected visual medium. In this process, based on prompt statements, the device automatically adjusts the tone of the video and sets background music according to instructions such as, "Edit the video of a scene of someone enjoying a walk in the park, add relaxing music at the moment the user smiles, and add a bright tone to the scenery."

[0861] Step 5:

[0862] The terminal presents the generated visual medium to the user through a user interface and accepts simple adjustments. The input is the generated visual medium, and the output is the medium adjusted based on the user's feedback. The terminal displays the video as a preview and provides an interface that allows the user to change scenes and fine-tune effects as needed.

[0863] Step 6:

[0864] The server automatically publishes the finalized visual media to the online platform via a posting device. The input is the adjusted visual media, and the output is the published video content. In this process, the media is uploaded to a pre-configured platform account, allowing users to easily share the content through social media and other means.

[0865] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0866] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0867] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0868] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0869] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0870] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0871] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0872] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0873] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0874] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0875] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0876] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0877] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0878] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0879] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0880] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0881] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0882] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0883] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0884] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0885] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0886] The following is further disclosed regarding the embodiments described above.

[0887] (Claim 1)

[0888] A means by which a recording device automatically captures video based on conditions,

[0889] A means for converting video data captured by a compression means into a predetermined format,

[0890] A means by which a transmitting device uploads compressed video data to a central processing unit via a network,

[0891] The analysis device uses AI technology to analyze video data and extract important scenes,

[0892] A means for generating a visual medium by editing video data extracted using AI technology,

[0893] A means of presenting a generated visual medium to the user and accepting adjustments,

[0894] A means of publishing a visual medium that has been adjusted for posting to an online platform,

[0895] ...

[0896] A system that includes this.

[0897] (Claim 2)

[0898] The system according to claim 1, further comprising means for an analysis device to analyze audio data and generate text through natural language processing.

[0899] (Claim 3)

[0900] The system according to claim 1, further comprising means for automatically applying mosaic processing based on personal identification information to a visual medium generated by a generating device.

[0901] "Example 1"

[0902] (Claim 1)

[0903] A means by which a recording device automatically acquires video based on conditions,

[0904] A means for converting video information acquired by a compression device into a predetermined format,

[0905] A means for a transmitting device to transmit compressed video information to a central control unit via a communication network,

[0906] The analysis device uses artificial intelligence technology to analyze video information and select specific scenes,

[0907] A means for generating a visual information medium by editing selected video information using artificial intelligence technology,

[0908] A means of presenting a user interface to the user as a generated visual information medium and accepting modifications,

[0909] A means for a posting device to transmit modified visual information media to an electronic platform,

[0910] A means that enables intuitive modification of the user interface using touch operation,

[0911] ...

[0912] A system that includes this.

[0913] (Claim 2)

[0914] The system according to claim 1, further comprising means for analyzing audio information and generating textual information through natural language processing.

[0915] (Claim 3)

[0916] The system according to claim 1, further comprising means for automatically applying visual correction processing based on personally identifiable information to a visual information medium generated by a generating device.

[0917] "Application Example 1"

[0918] (Claim 1)

[0919] A means by which a recording device automatically captures video based on conditions,

[0920] A means for converting video data captured by a compression means into a predetermined format,

[0921] A means by which a transmitting device uploads compressed video data to a central processing unit via a network,

[0922] The analysis device uses AI technology to analyze video data and extract important scenes,

[0923] A means for generating a visual medium by editing video data extracted using AI technology,

[0924] A means of presenting a generated visual medium to the user and accepting adjustments,

[0925] A means of publishing a visual medium that has been adjusted for posting to an online platform,

[0926] A means of delivering user-generated content based on visual media generated on smart devices,

[0927] ...

[0928] A system that includes this.

[0929] (Claim 2)

[0930] The system according to claim 1, further comprising means for an analysis device to analyze audio data and generate text through natural language processing.

[0931] (Claim 3)

[0932] The system according to claim 1, further comprising means for automatically applying mosaic processing based on personal identification information to a visual medium generated by a generating device.

[0933] "Example 2 of combining an emotion engine"

[0934] (Claim 1)

[0935] A means by which a recording device automatically acquires video based on conditions,

[0936] A means for converting video data acquired by a compression means into a predetermined format,

[0937] A means by which a transmitting device uploads compressed video data to an information processing device via a communication network,

[0938] The analysis device uses AI technology to analyze video data and identify important scenes, including the user's emotions.

[0939] A means for generating a visual medium by editing video data based on emotional data,

[0940] A means of presenting a user interface to the user as a generated visual medium and accepting simple adjustment operations,

[0941] A means of publishing a visual medium that has been adjusted by a posting device to an internet platform,

[0942] The means used to enable low-bandwidth communication as a compression means,

[0943] A system that includes this.

[0944] (Claim 2)

[0945] The system according to claim 1, further comprising means for analyzing speech data using an analysis device and generating spoken content as text using natural language processing.

[0946] (Claim 3)

[0947] The system according to claim 1, further comprising means for the generating device to automatically apply anonymization processing based on personal information protection to a visual medium.

[0948] "Application example 2 when combining with an emotional engine"

[0949] (Claim 1)

[0950] A means by which a recording device automatically captures video based on conditions,

[0951] A means for converting video data captured by a compression means into a predetermined format,

[0952] A means by which a transmitting device uploads compressed video data to a central processing unit via a network,

[0953] The analysis device uses AI technology to analyze video data and extract important scenes based on the user's emotions,

[0954] A means for generating a visual medium by editing video data extracted using AI technology based on an emotion engine,

[0955] A means of presenting a generated visual medium to the user and accepting adjustments,

[0956] A means of publishing a visual medium that has been adjusted for posting to an online platform,

[0957] A means to automatically optimize the tone and background sound of the video based on the results of user emotion analysis,

[0958] A system that includes this.

[0959] (Claim 2)

[0960] The system according to claim 1, characterized in that the analysis device includes means for analyzing audio data and generating text through natural language processing, and further includes means for setting importance based on the user's eye-tracking information.

[0961] (Claim 3)

[0962] The system according to claim 1, further comprising means for automatically applying mosaic processing based on personal identification information to a visual medium generated by a generating device, and further comprising means for automatically generating prompt sentences to be given to a generating AI model. [Explanation of symbols]

[0963] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means by which a recording device automatically captures video based on conditions, A means for converting video data captured by a compression means into a predetermined format, A means by which a transmitting device uploads compressed video data to a central processing unit via a network, The analysis device uses AI technology to analyze video data and extract important scenes, A means for generating a visual medium by editing video data extracted using AI technology, A means of presenting a generated visual medium to the user and accepting adjustments, A means of publishing a visual medium that has been adjusted for posting to an online platform, A system that includes this.

2. The system according to claim 1, further comprising means for an analysis device to analyze audio data and generate text through natural language processing.

3. The system according to claim 1, further comprising means for automatically applying mosaic processing based on personal identification information to a visual medium generated by a generating device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A