System

A system using real-time video capture, facial recognition, and automatic editing generates high-quality event videos, reducing the effort required for parents and enhancing parent-child interaction.

JP2026023928APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126249
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Filming and editing videos of children at events, such as sports days and school recitals, is tedious for parents, causing them to miss their child's facial expressions and touching moments, and requires significant time and effort, detracting from quality interaction.

Method used

A system using a high-resolution image capture device to record the entire event, real-time video transmission to a server, facial recognition to identify the subject, automatic clip generation and editing, and uploading to cloud storage, providing an intuitive user interface for access.

Benefits of technology

Eliminates the need for tedious filming and editing, allowing parents to spend more time with their children and easily record high-quality memories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023928000001_ABST
    Figure 2026023928000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for acquiring a video of an entire event using a high-resolution image capturing device; means for transmitting the video to a server in real time; means for identifying a subject using a facial recognition algorithm on the server; means for tracking the subject in the video; means for automatically generating a clip focused on the subject; means for editing the clip to generate an individual video file; means for uploading the generated video file to a cloud storage; and means for providing a user interface capable of accessing the video file.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] This system aims to solve the problem of the time and effort required for filming and editing videos of children at events (such as sports days and school recitals) for families with children. Filming the entire event, selecting and editing footage of a specific child, is particularly tedious, leading to parents missing their child's facial expressions and touching moments. Therefore, there is a need to increase the amount of time parents can spend with their children and realize more intimate parent-child interactions. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides the following means:

[0006] means for capturing video of the entire event using a high resolution image capture device;

[0007] means for transmitting the video to a server in real time;

[0008] a means for identifying the subject using a facial recognition algorithm on the server;

[0009] means for tracking the subject within the video;

[0010] means for automatically generating clips focused on the subject;

[0011] means for editing said clips to generate individual video files;

[0012] A means for uploading the generated video file to cloud storage;

[0013] means for providing a user interface that allows access to the video file;

[0014] The present invention provides a system including:

[0015] This eliminates the need for tedious filming and editing tasks during events, allowing parents to spend more time with their children and enjoying the event. Furthermore, the system automatically generates high-quality footage, making it easy and efficient to record memories.

[0016] A "high resolution image capture device" is a high quality camera device that captures the entire event in detail.

[0017] "Means for transmitting video to a server in real time" refers to communication technology and protocols for immediately transferring captured video data to a server.

[0018] The "server" is a central processing unit that processes received video data and stores and manages it in a database.

[0019] A "facial recognition algorithm" is a machine learning technology and software that recognizes specific faces in video and identifies that person.

[0020] "Means for tracking a person of interest within a video" refers to software technology for continuously tracking the location of an identified person within a video frame.

[0021] The "means for automatically generating clips" refers to algorithms and software for extracting specific video portions based on the tracking results and saving them as independent video clips.

[0022] "Means for generating a video file" refers to a video editing technique for editing multiple video clips and saving them as one continuous video.

[0023] "Means for uploading video files to cloud storage" refers to communication technologies and protocols for storing generated video files on a remote server on the Internet.

[0024] "User Interface" means the software operating screen that allows users to access the system and view and download the generated video files. [Brief explanation of the drawings]

[0025] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0026] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0027] First, the terms used in the following description will be explained.

[0028] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0029] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0030] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0031] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0033] [First embodiment]

[0034] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0035] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0036] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0037] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0038] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0040] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0041] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0042] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0043] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0044] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0045] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0046] The present invention is a system for automatically generating personalized videos for families with children, allowing parents to easily record their children's memories without the need to film and edit events. This system is implemented as follows.

[0047] System configuration

[0048] 1. High-resolution imaging equipment

[0049] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[0050] 2. Real-time video transmission method

[0051] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[0052] 3. Facial Recognition and Video Analysis

[0053] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0054] 4. How to track the subject

[0055] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0056] 5. Automatic Clip Generation Method

[0057] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0058] 6. Video file generation method

[0059] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[0060] 7. Cloud Storage Distribution Methods

[0061] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0062] 8. User Interface

[0063] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0064] Specific examples

[0065] Sports day case

[0066] 1. Event Settings

[0067] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0068] 2. Start shooting

[0069] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0070] 3. Video data transmission

[0071] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0072] 4. Clip Creation and Video Editing

[0073] The server automatically extracts and edits the footage of the child in question into a clip. For example, a clip of the child running is clipped and edited to include all the necessary scenes.

[0074] 5. Distribution of video files

[0075] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0076] This system allows parents to easily record and enjoy their children's growth and memories without the hassle of filming and editing during events, increasing the amount of time parents can spend cheering and interacting with their children at the venue and creating a more fulfilling parent-child experience.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[0080] Step 2:

[0081] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[0082] Step 3:

[0083] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[0084] Step 4:

[0085] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[0086] Step 5:

[0087] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0088] Step 6:

[0089] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[0090] Step 7:

[0091] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[0092] Step 8:

[0093] The server then edits the multiple video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[0094] Step 9:

[0095] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[0096] Step 10:

[0097] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[0098] This series of processes eliminates the need for the user to perform tedious shooting and editing tasks during the event, allowing them to have more time to enjoy their children's growth and memories.

[0099] Example 1

[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0101] Nowadays, parents spend a lot of time and effort recording their children's growth and important events. Filming and editing events is particularly tedious work, requiring parents to spend a lot of time on the task, which can result in parents being unable to fully enjoy the event itself. Furthermore, the quality of the footage recorded varies depending on the cameraman's skill and the equipment used. Therefore, there is a need for a system that allows parents to easily obtain high-quality footage.

[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0103] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video in real time to a data processing device, means for identifying a subject using a facial recognition algorithm on the data processing device, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for automatically zooming and frame correction for the subject, and means for pre-registering facial images of the subject. This allows parents to easily record and enjoy high-quality video of their children's growth and activities during the event without the effort of filming or editing.

[0104] "High-resolution image capturing device" refers to any imaging device that has a high pixel count and can capture detailed, clear images.

[0105] "Means for transmitting video in real time" refers to communication technology and devices for transferring captured video data to a server or data processing device immediately without delay.

[0106] A "facial recognition algorithm" refers to software or computational methods used to identify faces in images and detect specific individuals.

[0107] "Means for identifying subjects" refers to functions and technologies for recognizing and identifying individual people within video footage.

[0108] "Means for tracking a subject within a video stream" refers to the technology or computational methods used to continuously track the location of an identified person as they move within the video stream.

[0109] "Means for automatically generating clips" refers to functions and technologies for cutting out footage showing specific people and automatically editing it into short video segments.

[0110] "Means for generating individual video files" refers to the functions and technologies for editing and integrating multiple video clips into a continuous video file.

[0111] "Means for uploading to cloud storage" refers to the functionality and technology for storing the generated video files in a cloud-based storage device via the Internet.

[0112] "Means for providing a user interface" refers to display screens and operating means that make it easier for users to operate systems and services.

[0113] "Means for automatic zoom and frame correction" refers to functions and technologies for zooming in on a specific subject in a video or automatically adjusting the composition of the video.

[0114] "Means for pre-registering facial images" refers to functions and technologies that allow the system to load and store facial image data of the subject in advance.

[0115] MODE FOR CARRYING OUT THE INVENTION

[0116] This invention is an automatic individual video generation system designed for families with children, which reduces the effort required for parents to film and edit events and allows them to easily record their children's growth and memories. This system is implemented using the following configuration and process.

[0117] System configuration

[0118] 1. High-resolution imaging equipment

[0119] Users install a high-resolution camera at the event venue. This camera takes still pictures and captures the entire event in detail. This device can be, for example, a 4K video camera or a high-performance digital SLR camera.

[0120] 2. Real-time video transmission

[0121] The device transmits the captured video data to a server in real time. This process uses an internet connection, so the data is sent to the server instantly. Specifically, a smartphone or tablet receives the data from the camera and sends it to the server via Wi-Fi or mobile network.

[0122] 3. Facial Recognition and Video Analysis

[0123] The server analyzes the received video data and uses AI facial recognition algorithms to identify the specific child's face, using facial recognition software such as Amazon Rekognition or OpenCV.

[0124] 4. Subject Tracking

[0125] The server tracks the recognized child's face in the video and records its position in each frame, allowing it to track the child's position even if the child moves within the video.

[0126] 5. Automatic generation of video clips

[0127] The server automatically extracts the footage containing the subject and generates individual video clips using video editing software such as FFmpeg, and automatically applies zoom and frame correction to create visually pleasing clips.

[0128] 6. Generating video files

[0129] The server then assembles the edited video clips into a series of video files, a process that organizes the clips chronologically to create a coherent video.

[0130] 7. Upload to cloud storage

[0131] The server then uploads the generated video files to cloud storage, typically a cloud-based storage service such as Google Drive or Dropbox.

[0132] 8. Providing a User Interface

[0133] Users access video files stored in cloud storage through a dedicated application, which is intuitive and easy to use, designed to allow users to easily find, download, and watch videos.

[0134] Specific examples

[0135] In the case of sports day

[0136] 1. Event setup: The user sets up the "Sports Day" event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0137] 2. Start shooting: The user sets up a high-resolution camera at the sports day venue and starts shooting from a position that covers the entire venue.

[0138] 3. Real-time video transmission: The device transmits video data to the server in real time. The server analyzes the received video and performs facial recognition and tracking of the subject.

[0139] 4. Clip generation and video editing: The server automatically extracts and edits the video clips containing the target children. For example, it clips the footage of the children running and edits it to include all the necessary scenes.

[0140] 5. Video file distribution: The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0141] Prompt Sentence Examples

[0142] markdown

[0143] Describe a system that automatically edits and saves footage of children's sports events to cloud storage. The user installs a high-resolution camera and configures the event using a dedicated app. The device transmits the footage in real time to a server, which uses AI facial recognition to identify the children and automatically edits the footage. The final video file is uploaded to cloud storage and can be accessed by the user.

[0144] This system allows parents to record and enjoy high-quality footage of their children's growth and memories without the effort of filming and editing during events. By increasing the amount of time parents spend cheering and interacting with their children at the venue, parents can spend more fulfilling time with their children.

[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0146] Step 1:

[0147] The user launches a dedicated application and sets up the event. First, they move to the settings screen within the application and enter the event name (e.g., "Sports Day") and date. Next, they upload a facial image of the child in question. This facial image will be used for facial recognition later. The event information and facial image set by the user are entered. The entered data is saved within the application.

[0148] Step 2:

[0149] Users set up high-resolution cameras at the event venue. They position the cameras so that they can cover the entire event, and adjust the angle and height as needed. When they're ready to start recording, they press the camera's record button to begin recording. The cameras generate high-quality video data in real time, allowing the entire event to be recorded in detail.

[0150] Step 3:

[0151] The terminal (device connected to the camera) compresses the captured video data in real time and sends it to a server via the Internet. Video data is input and output in a compressed form. The video data sent from the terminal is received by the server. This process achieves efficient data transfer.

[0152] Step 4:

[0153] The server analyzes the received video data and uses an AI facial recognition algorithm (such as Amazon Rekognition or OpenCV) to match it with previously uploaded facial images. The facial image is input and the face of a specific child is recognized. The server identifies the target person based on this processing. The location of the specific child within the video is confirmed through facial recognition.

[0154] Step 5:

[0155] The server tracks the recognized subject's face in the video and records their location in each frame. The location information obtained through facial recognition is input, and tracking data including the subject's location is output. The server can use this tracking data to track the location of a specific child over time.

[0156] Step 6:

[0157] The server automatically extracts the portion of the video that shows the subject and generates a separate video clip. This process is performed using video editing software (e.g., FFmpeg). Tracking data is input and a video clip with the subject at the center is output. The server then applies zoom and frame correction to this clip to make it more visually appealing.

[0158] Step 7:

[0159] The server chronologically organizes multiple video clips into a series of video files. Edited clips are input and a coherent video file is output. The server encodes this video file and converts it to the optimal format (e.g., MP4).

[0160] Step 8:

[0161] The server uploads the generated video file to the cloud storage specified by the user (e.g., Google Drive or Dropbox). The completed video file is input and saved in the cloud storage. The upload process sends the file over the Internet and stores it securely in the cloud.

[0162] Step 9:

[0163] Users access cloud storage through a dedicated application to view uploaded video files. Users can search for video files and download or stream them. The application provides an intuitive interface, making it easy to find the video you want.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] In conventional virtual shopping, parents must manually record and then edit videos to record their children's real-time reactions. This process requires time and effort for parents, hindering an intuitive and natural shopping experience. Furthermore, manual filming and editing increases the risk of missing important moments. Therefore, there is a need for a system that can automatically capture and record children's reactions and fun moments.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, and means for automatically capturing and recording a child's reactions and moments of enjoyment during virtual shopping. This eliminates the need for parents to manually shoot and edit videos, and allows parents to easily record important moments during virtual shopping without missing them.

[0169] A "high resolution image capture device" is a camera device capable of capturing video at high resolution.

[0170] "Means for transmitting to the server in real time" refers to technology that transfers video data to the server immediately without delay.

[0171] "Means for identifying a subject using a facial recognition algorithm on a server" refers to a method in which a server uses facial recognition technology to identify a specific person.

[0172] "Means for tracking a subject in a video" refers to a technology that continuously tracks the position and movements of a specific person in a video.

[0173] "Means for automatically generating clips" refers to a technology that automatically cuts out specific parts of video and edits them into short video segments.

[0174] The "means for generating a video file" refers to a method for integrating multiple video clips and editing them into a single continuous video file.

[0175] "Means for uploading to cloud storage" refers to a technology that stores the generated video files on a remote server and makes them accessible via the Internet.

[0176] "Means for providing a user interface" refers to operation screens and applications that allow users to interact with the system intuitively.

[0177] "Virtual shopping" is the act of selecting and purchasing products in a virtual space using the Internet.

[0178] "A means for automatically capturing and recording children's reactions and moments of enjoyment" is a technology that detects children's facial expressions and behavior in real time and automatically records those moments as video.

[0179] This invention is a system that automatically captures and records moments that parents enjoy with their children during virtual shopping, allowing parents to preserve memories without any hassle.

[0180] System configuration

[0181] 1. High-resolution imaging equipment

[0182] Users place a high-resolution camera in the virtual shopping environment to capture detailed, high-resolution footage of the entire shopping experience. The camera remains stationary and captures footage of the shopping experience.

[0183] 2. Real-time video transmission method

[0184] The device (smartphone or head-mounted display) transmits the captured video data to a server in real time using an internet connection, and video data captured while shopping is sent to the server instantly.

[0185] 3. Facial Recognition and Video Analysis

[0186] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0187] 4. How to track the subject

[0188] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0189] 5. Automatic Clip Generation Method

[0190] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0191] 6. Video file generation method

[0192] The server compiles the video clips into a series of video files, which are organized chronologically to create a coherent video of the entire shopping experience.

[0193] 7. Cloud Storage Distribution Methods

[0194] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0195] 8. User Interface

[0196] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0197] Hardware and software used

[0198] Hardware

[0199] Webcam

[0200] Smartphone

[0201] head-mounted display

[0202] software

[0203] OpenCV (image processing library)

[0204] dlib (face recognition algorithm)

[0205] moviepy (video editing library)

[0206] Specific application examples

[0207] The case for virtual shopping

[0208] 1. Pre-registration of facial images

[0209] Users upload a photo of their child's face in advance through a dedicated application.

[0210] 2. Start shooting

[0211] Users place high-resolution cameras in the virtual shopping environment and capture footage in real time.

[0212] 3. Video data transmission

[0213] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0214] 4. Clip Creation and Video Editing

[0215] The server automatically cuts out clips of footage in which the child appears, and edits them together, for example, when the child shows interest in a product, and edits the clips to include all the important moments.

[0216] 5. Distribution of video files

[0217] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0218] Prompt Sentence Examples

[0219] "Develop an application that automatically captures a child's smile or surprised expression while they are shopping, and creates a video of their memories that can be viewed later."

[0220] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0221] Step 1:

[0222] Users first upload a facial image of their child through a dedicated application. In this step, users register their child's facial image data in the system using a smartphone or computer. The system receives this image data as input and generates and stores facial feature data for use in the facial recognition algorithm.

[0223] Step 2:

[0224] The user installs a high-resolution camera in the virtual shopping environment and captures video in real time. The video captured by the camera is positioned to cover the entire view of the shop. The camera captures video data in real time and transmits it to the terminal. The terminal then transmits the received video data to the server in real time.

[0225] Step 3:

[0226] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face. In this step, the server receives real-time video data as input and runs the facial recognition algorithm. The algorithm compares the face in the real-time video with pre-registered facial feature data, and if there is a match, the child's face is identified.

[0227] Step 4:

[0228] The server tracks the recognized subject's face in the video and records its position in each frame. In this step, the server continues to record the position of the identified face in each frame of the video. For each frame received as input, it provides the face position information as output.

[0229] Step 5:

[0230] The server automatically extracts the video portion where the subject appears and generates a separate video clip. In this step, the server analyzes the video based on the facial position information, extracts the portion where the subject appears, and creates a clip. Zoom and frame correction are also applied automatically. The input is the tracked video frame and position information, and the output is a visually pleasing video clip.

[0231] Step 6:

[0232] The server edits multiple video clips and compiles them into a series of video files. In this step, the server integrates the generated video clips to create a time-coherent video file. It takes individual video clips as input, edits them, and generates a continuous video file as output.

[0233] Step 7:

[0234] The server uploads the generated video file to the cloud storage. In this step, the server saves the completed video file to the cloud storage via the Internet. It receives the completed video file as input and provides the cloud storage URL or access information as output.

[0235] Step 8:

[0236] Users access, download, and watch video files stored in cloud storage through a dedicated application. In this step, users access cloud storage using a dedicated user interface to view and save video files. Cloud storage access information is used as input, and video files are made available for viewing as output.

[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0238] The present invention is a system that automatically generates individual videos for families with children and then uses an emotion engine to edit the videos based on the user's emotions. This system eliminates the need for parents to film and edit events, allowing them to easily record their children's memories and highlight moving moments. Specific embodiments of the system are described below.

[0239] System configuration

[0240] 1. High-resolution imaging equipment

[0241] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[0242] 2. Real-time video transmission method

[0243] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[0244] 3. Facial Recognition and Video Analysis

[0245] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0246] 4. How to track the subject

[0247] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0248] 5. Automatic Clip Generation Method

[0249] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0250] 6. Video file generation method

[0251] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[0252] 7. Cloud Storage Distribution Methods

[0253] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0254] 8. User Interface

[0255] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0256] Incorporating an emotion engine

[0257] 1. Emotion Engine

[0258] The server also has an emotion engine that recognizes the user's emotions from the video, analyzing emotions such as smile, sadness, and surprise in real time to obtain emotional data during the event.

[0259] 2. Emotional editing

[0260] The server then edits the video clips based on the recognized emotions, for example, prioritizing scenes with many smiling faces as highlight clips, or editing to emphasize moving scenes.

[0261] Specific examples

[0262] Emotion Recognition in the Case of Sports Day

[0263] 1. Event Settings

[0264] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0265] 2. Start shooting

[0266] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0267] 3. Video data transmission

[0268] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0269] 4. Emotion analysis

[0270] The server also uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[0271] 5. Clip Creation and Video Editing

[0272] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[0273] 6. Distribution of video files

[0274] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0275] This system allows users to record their children's growth and memories in a richer way, without the hassle of filming and editing during events, and enjoy footage that highlights moving moments. This increases the amount of time parents and children can spend together cheering and interacting with each other at the venue, allowing for more fulfilling parent-child time.

[0276] The processing flow will be explained below.

[0277] Step 1:

[0278] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[0279] Step 2:

[0280] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[0281] Step 3:

[0282] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[0283] Step 4:

[0284] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[0285] Step 5:

[0286] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0287] Step 6:

[0288] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[0289] Step 7:

[0290] The server also uses an emotion engine to analyze the emotions of the subjects in the video and the viewers in real time, detecting, for example, the smile or other emotions of a child crossing the finish line in a race.

[0291] Step 8:

[0292] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[0293] Step 9:

[0294] The server determines the importance of each clip based on the emotional data recognized by the emotion engine, and prioritizes highlights, such as scenes of smiling faces or cheering.

[0295] Step 10:

[0296] The server then edits the selected video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[0297] Step 11:

[0298] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[0299] Step 12:

[0300] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[0301] This series of processes eliminates the need for users to perform tedious filming and editing tasks during events, allowing them to spend more time enjoying their children's growth and memories. Furthermore, the emotion engine makes it possible to create videos that emphasize moving moments.

[0302] Example 2

[0303] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0304] For families with children, the time and effort required to film and edit events is a challenge. In particular, if parents are too focused on filming during an event, they lose the time they need to cheer on and interact with their children in real time. Furthermore, editing the footage to appropriately highlight moving moments and records of children's growth is difficult.

[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0306] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject in the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for recognizing emotions in the video using an emotion engine, and means for editing the video based on the recognized emotions. This allows users to easily create videos that highlight their children's growth and moving moments without the hassle of filming and editing the event.

[0307] "High-resolution image capture device" refers to a capture device capable of capturing detailed footage of the entire event.

[0308] "Means for transmitting to a server in real time" refers to a communication means for immediately transferring captured video data to a server.

[0309] A "facial recognition algorithm" is software or a program that identifies the faces of people in a video and authenticates a specific individual.

[0310] "Subject" refers to an individual identified by a facial recognition algorithm.

[0311] A "tracking means" is a system for continuously recording the location of a subject recognized in the video.

[0312] The "means for automatically generating clips" is a function that automatically cuts out the portion of the video in which the subject appears.

[0313] The "means for generating a video file" is a system that combines multiple video clips and edits them into a single video file.

[0314] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.

[0315] "User interface" refers to the screen and input means used by users to operate software or a system.

[0316] An "emotion engine" is software or an algorithm that analyzes the emotions of people in a video and recognizes their emotional state.

[0317] "Means for editing video" refers to a function for processing captured video data and emphasizing or changing content based on specific scenes or emotions.

[0318] This system automatically generates individual videos for families with children and uses an emotion engine to edit the videos based on the user's emotions. This saves parents the trouble of filming and editing events, allowing them to easily record their children's memories and highlight moving moments.

[0319] This system is implemented using the following hardware and software.

[0320] High-resolution image capture device: The user installs a high-resolution camera at the event venue to capture the entire event. For example, the Sony Alpha series is used.

[0321] Real-time video transmission method: The device (e.g., laptop or Raspberry Pi) transmits the video data received from the camera to the server in real time. The communication method is a high-speed Internet connection (Wi-Fi or wired LAN).

[0322] Facial recognition algorithm: The server uses an AI facial recognition algorithm such as Amazon Rekognition to recognize the child's face from the received video data and identify the child by comparing it with previously uploaded facial image data.

[0323] Subject tracking method: The server continuously tracks the location information of the recognized subject within the video. It records the position in each frame, so it can follow the subject even if they move within the video.

[0324] Automatic clip generation: The server automatically extracts the video portion in which the recognized subject appears and generates individual video clips. The clips are then edited to be visually appealing, with automatic zooming and frame correction.

[0325] Video file generator: The server edits multiple video clips and compiles them into a series of video files, creating a consistent video of the entire event.

[0326] Cloud storage distribution method: The server uploads the completed video file to cloud storage (e.g., Google Drive). The data is saved in the cloud storage specified by the user.

[0327] User interface: Users can access cloud storage through a dedicated app (iOS or Android app) and download or watch video files.

[0328] Emotion engine: The server uses Amazon Rekognition and Microsoft Azure Emotion API to analyze the user's emotions in real time. It recognizes emotions such as smiles and surprise and records data during the event.

[0329] Emotion-based editing: The server edits video clips based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights and emphasizes moving scenes.

[0330] Example: Emotion recognition in the case of athletic meet

[0331] 1. Event Settings

[0332] Users set up a "Sports Day" event using a dedicated app, and upload facial images of the children involved in the event in advance.

[0333] 2. Start shooting

[0334] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0335] 3. Video data transmission

[0336] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0337] 4. Emotion analysis

[0338] The server uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[0339] 5. Clip Creation and Video Editing

[0340] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[0341] 6. Distribution of video files

[0342] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated app.

[0343] Example prompts for generative AI models:

[0344] Perform emotion analysis on video data from a sports day and generate a video clip that emphasizes touching moments. The facial images of the target children and the actual video data can be found at the following links: (Facial image URL), (Video data URL). Edit the scenes in which smiling faces are detected as highlights.

[0345] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0346] Step 1:

[0347] The user sets up a high-resolution camera at the event venue, fixes the camera in a position where it can capture the entire event in detail, and starts recording, inputting video data of the entire event.

[0348] Step 2:

[0349] The device (laptop or Raspberry Pi) transmits the video data acquired from the camera to the server in real time. The communication method is a high-speed internet connection. The input in this step is the video data from the camera, and the output is the real-time video data transmission to the server.

[0350] Step 3:

[0351] The server analyzes the received video data using AI facial recognition algorithms such as Amazon Rekognition. It compares the data with previously uploaded facial image data and recognizes the child's face. The input for this step is real-time video data and facial image data, and the output is the location information of the recognized face.

[0352] Step 4:

[0353] The server tracks the recognized face in the video and keeps recording the subject's position information for each frame. The input is the position information of the recognized face, and the output is the subject's position data for each frame.

[0354] Step 5:

[0355] The server automatically extracts the portion of the video in which the recognized and tracked subject appears, generating individual video clips. The clips are then edited for visual clarity with automatic zoom and frame correction. The inputs are the subject's position data and video data, and the output is the video clip.

[0356] Step 6:

[0357] The server then edits the generated video clips into a series of video files. The clips are organized in chronological order to create a seamless, coherent video. The input is a series of video clips, and the output is an edited video file.

[0358] Step 7:

[0359] The server uploads the completed video file to the specified cloud storage, such as Google Drive or Amazon S3. The input is the edited video file, and the output is a notification that the upload to the cloud storage is complete.

[0360] Step 8:

[0361] Users use a dedicated app to access cloud storage and download or view video files. The interface is intuitive and easy to use. The input is the cloud storage URL, and the output is a video file that can be downloaded and viewed.

[0362] Step 9:

[0363] The server uses an emotion engine (such as Amazon Rekognition or Microsoft Azure Emotion API) to analyze the user's emotions in the video in real time. It recognizes emotions such as smiles and surprise and records emotional data during the event. The input is the video data, and the output is the emotion recognition results.

[0364] Step 10:

[0365] The server then edits the video clip based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights, and further emphasizes moving scenes. The input is the emotion recognition results and the video clip, and the output is an edited video based on the emotions.

[0366] (Application example 2)

[0367] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0368] In modern factory environments, worker safety management is a critical issue. Conventional monitoring methods have difficulty monitoring in real time whether workers are working in an appropriate state of health or under emotional stress. Furthermore, dangerous situations frequently occur at worksites due to the large number of machines in operation. Conventional safety monitoring systems rely on simple surveillance cameras, which are unable to analyze workers' emotions or fatigue levels. This makes it difficult to adequately ensure safety in the actual work environment. Therefore, there is a need for a system that can analyze workers' emotional state and fatigue levels in real time and respond quickly.

[0369] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video using a high-resolution image capture device, means for transmitting the video in real time, means for identifying a target person using a facial recognition algorithm, means for tracking the target person in the video, means for automatically generating clips, means for editing the clips to generate a video file, means for uploading the video file to cloud storage, means for providing a user interface for accessing the video file, means for analyzing emotions in the video in real time using an emotion engine and issuing a safety warning, and means for transmitting the analysis data in real time, and means for storing the data in cloud storage. This makes it possible to monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[0370] A "high resolution imaging device" is an imaging device for capturing highly detailed videos and images.

[0371] "Means for transmitting to the server in real time" refers to the communication technology for instantly sending captured video data to the server.

[0372] A "face recognition algorithm" is a computer program that identifies faces in video and compares them with pre-registered facial information.

[0373] "Means for tracking a subject within a video" refers to a technology that tracks the position of a specific person even if they move within the video.

[0374] "Means for automatically generating clips" refers to technology that automatically cuts out specific video segments and generates them as short video clips.

[0375] The "means for generating a video file" refers to the technology for editing clips and putting them together as a series of videos.

[0376] "Means for uploading to cloud storage" refers to a technology for storing the generated video file in a remote storage area on the Internet.

[0377] A "user interface" is the portion of software that provides the visual and operational elements through which a user interacts with a system.

[0378] An "emotion engine" is an algorithm and software that analyzes the emotions of people in a video and identifies their emotional state.

[0379] "Means for issuing safety warnings" refers to technology for issuing warnings when dangerous situations are detected as a result of emotion analysis.

[0380] "Data transmission means" refers to a communication technique for sending the analyzed data to another location.

[0381] "Means by which data is stored in cloud storage" refers to the technology used to store analyzed data and generated videos in the cloud.

[0382] The present invention relates to a worker safety monitoring system for a factory, which uses a high-resolution image capture device, a high-performance server, and cloud storage. The system analyzes the facial expressions and postures of workers working in the factory and can perform safety management in real time. Specific embodiments of the system are described below.

[0383] System configuration

[0384] 1. High-resolution imaging equipment

[0385] The user installs a high-resolution camera in the factory, which covers a wide area and captures detailed images of workers. For example, the Logitech C920 is used as this camera.

[0386] 2. Real-time video transmission method

[0387] The device transmits the captured video data in real time to a server using an internet connection, where the video data is processed immediately.

[0388] 3. Facial Recognition Algorithm

[0389] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the worker's face. At this time, it compares it with pre-registered facial images to identify the specific worker. OpenCV is used for facial recognition.

[0390] 4. How to track the subject

[0391] The server tracks the recognized worker's face in the video and records the position information in each frame, making it possible to track the target's position even if the target moves within the video.

[0392] 5. Emotion analysis

[0393] The server uses an emotion engine to analyze the emotions of workers in the video in real time. For example, it analyzes emotions such as smiles, sadness, and surprise, and generates data useful for safety management. An AI model using TensorFlow / Keras is used for emotion analysis.

[0394] 6. Issuance of safety warnings

[0395] Based on the analysis results of the emotion engine, the server quickly issues safety warnings, including visual alerts and audio notifications, if a worker is in a dangerous or overly fatigued state.

[0396] 7. Data transmission and storage in cloud storage

[0397] The server uploads the analyzed data and generated video clips to cloud storage and also transmits the analyzed data to other devices in real time, allowing factory managers to monitor the safety status of workers even from remote locations.

[0398] 8. User Interface

[0399] Users can access cloud storage and check saved data and video clips using a dedicated application. The intuitive user interface allows users to quickly obtain the information they need.

[0400] Specific examples

[0401] To implement a system to monitor the safety of workers in a factory, high-resolution cameras are installed in each work area. The cameras capture the facial expressions and postures of workers in real time and send the footage to a server. The server then uses a facial recognition algorithm and an emotion engine to identify the worker and perform emotion analysis. Based on the analysis results, a safety warning is issued immediately if the worker is in danger. Furthermore, the real-time analysis data is stored in cloud storage and can be accessed by managers through a dedicated application.

[0402] Example prompts for generative AI models

[0403] Design a system that uses facial recognition and emotion analysis to monitor the safety of workers in a factory. A high-resolution camera captures images of workers working and analyzes them in real time using an emotion engine. If a worker is in danger or overly fatigued, an alert will be issued and the data will be stored in the cloud. Please provide a concrete code example.

[0404] As a result, a system can be provided that can monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[0405] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0406] Step 1:

[0407] Users install high-resolution cameras in the work area of ​​their factories, which capture detailed images of workers' faces and the work they are doing.

[0408] Input: spatial information (camera installation position), high-resolution camera

[0409] Output: Real-time video data

[0410] Step 2:

[0411] The device transmits the captured real-time video data to a server via the Internet, where streaming technology is used to efficiently transmit large amounts of data.

[0412] Input: Real-time video data

[0413] Output: Video data sent to the server

[0414] Step 3:

[0415] The server uses OpenCV to analyze the received video data and identifies the worker's face using a facial recognition algorithm, which then compares it with pre-registered facial information to identify each worker.

[0416] Input: Transmitted video data, pre-registered face data

[0417] Output: Video data including identified workers (with location information)

[0418] Step 4:

[0419] The server tracks the face of the identified worker and records their position in each frame, allowing the worker to be tracked in real time even if they move within the camera's field of view.

[0420] Input: Video data containing identified workers

[0421] Output: Video data with location information of tracked workers

[0422] Step 5:

[0423] The server automatically generates a clip of the worker's face, zooming and adjusting the frame as needed, using an AI model to generate the optimal clip.

[0424] Input: Location-based video data of tracked workers

[0425] Output: Automatically generated clip video data

[0426] Step 6:

[0427] The server uses an emotion engine to analyze the emotions of the workers in the video clips, specifically identifying emotional states such as smiling or surprised using TensorFlow / Keras.

[0428] Input: Clip video data

[0429] Output: Parsed emotion data (with emotion labels)

[0430] Step 7:

[0431] The server then issues safety alerts based on the analyzed emotion data, for example, visual and audio alerts if a worker is in a dangerous emotional state (anger or sadness).

[0432] Input: Emotion-labeled analysis data

[0433] Output: Safety warning (alert)

[0434] Step 8:

[0435] The server uploads the analyzed data and generated video clips to cloud storage, and also transmits the analyzed data in real time to a remote administrator terminal.

[0436] Input: Analysis data, clip video data

[0437] Output: Data stored in the cloud, analysis data sent to the administrator's terminal

[0438] Step 9:

[0439] Users (administrators) can access cloud storage and check stored data and video clips using a dedicated application. The application has an intuitive interface, allowing users to quickly obtain the information they need.

[0440] Input: Data stored in the cloud (video clips, analysis data)

[0441] Output: Monitoring information available to administrators (interface operations)

[0442] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0443] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0444] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0445] [Second embodiment]

[0446] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0447] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0448] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0449] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0450] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0451] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0452] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0453] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0454] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0455] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0456] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0457] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0458] The present invention is a system for automatically generating personalized videos for families with children, allowing parents to easily record their children's memories without the need to film and edit events. This system is implemented as follows.

[0459] System configuration

[0460] 1. High-resolution imaging equipment

[0461] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[0462] 2. Real-time video transmission method

[0463] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[0464] 3. Facial Recognition and Video Analysis

[0465] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0466] 4. How to track the subject

[0467] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0468] 5. Automatic Clip Generation Method

[0469] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0470] 6. Video file generation method

[0471] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[0472] 7. Cloud Storage Distribution Methods

[0473] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0474] 8. User Interface

[0475] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0476] Specific examples

[0477] Sports day case

[0478] 1. Event Settings

[0479] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0480] 2. Start shooting

[0481] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0482] 3. Video data transmission

[0483] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0484] 4. Clip Creation and Video Editing

[0485] The server automatically extracts and edits the footage of the child in question into a clip. For example, a clip of the child running is clipped and edited to include all the necessary scenes.

[0486] 5. Distribution of video files

[0487] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0488] This system allows parents to easily record and enjoy their children's growth and memories without the hassle of filming and editing during events, increasing the amount of time parents can spend cheering and interacting with their children at the venue and creating a more fulfilling parent-child experience.

[0489] The processing flow will be explained below.

[0490] Step 1:

[0491] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[0492] Step 2:

[0493] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[0494] Step 3:

[0495] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[0496] Step 4:

[0497] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[0498] Step 5:

[0499] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0500] Step 6:

[0501] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[0502] Step 7:

[0503] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[0504] Step 8:

[0505] The server then edits the multiple video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[0506] Step 9:

[0507] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[0508] Step 10:

[0509] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[0510] This series of processes eliminates the need for the user to perform tedious shooting and editing tasks during the event, allowing them to have more time to enjoy their children's growth and memories.

[0511] Example 1

[0512] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0513] Nowadays, parents spend a lot of time and effort recording their children's growth and important events. Filming and editing events is particularly tedious work, requiring parents to spend a lot of time on the task, which can result in parents being unable to fully enjoy the event itself. Furthermore, the quality of the footage recorded varies depending on the cameraman's skill and the equipment used. Therefore, there is a need for a system that allows parents to easily obtain high-quality footage.

[0514] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0515] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video in real time to a data processing device, means for identifying a subject using a facial recognition algorithm on the data processing device, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for automatically zooming and frame correction for the subject, and means for pre-registering facial images of the subject. This allows parents to easily record and enjoy high-quality video of their children's growth and activities during the event without the effort of filming or editing.

[0516] "High-resolution image capturing device" refers to any imaging device that has a high pixel count and can capture detailed, clear images.

[0517] "Means for transmitting video in real time" refers to communication technology and devices for transferring captured video data to a server or data processing device immediately without delay.

[0518] A "facial recognition algorithm" refers to software or computational methods used to identify faces in images and detect specific individuals.

[0519] "Means for identifying subjects" refers to functions and technologies for recognizing and identifying individual people within video footage.

[0520] "Means for tracking a subject within a video stream" refers to the technology or computational methods used to continuously track the location of an identified person as they move within the video stream.

[0521] "Means for automatically generating clips" refers to functions and technologies for cutting out footage showing specific people and automatically editing it into short video segments.

[0522] "Means for generating individual video files" refers to the functions and technologies for editing and integrating multiple video clips into a continuous video file.

[0523] "Means for uploading to cloud storage" refers to the functionality and technology for storing the generated video files in a cloud-based storage device via the Internet.

[0524] "Means for providing a user interface" refers to display screens and operating means that make it easier for users to operate systems and services.

[0525] "Means for automatic zoom and frame correction" refers to functions and technologies for zooming in on a specific subject in a video or automatically adjusting the composition of the video.

[0526] "Means for pre-registering facial images" refers to functions and technologies that allow the system to load and store facial image data of the subject in advance.

[0527] MODE FOR CARRYING OUT THE INVENTION

[0528] This invention is an automatic individual video generation system designed for families with children, which reduces the effort required for parents to film and edit events and allows them to easily record their children's growth and memories. This system is implemented using the following configuration and process.

[0529] System configuration

[0530] 1. High-resolution imaging equipment

[0531] Users install a high-resolution camera at the event venue. This camera takes still pictures and captures the entire event in detail. This device can be, for example, a 4K video camera or a high-performance digital SLR camera.

[0532] 2. Real-time video transmission

[0533] The device transmits the captured video data to a server in real time. This process uses an internet connection, so the data is sent to the server instantly. Specifically, a smartphone or tablet receives the data from the camera and sends it to the server via Wi-Fi or mobile network.

[0534] 3. Facial Recognition and Video Analysis

[0535] The server analyzes the received video data and uses AI facial recognition algorithms to identify the specific child's face, using facial recognition software such as Amazon Rekognition or OpenCV.

[0536] 4. Subject Tracking

[0537] The server tracks the recognized child's face in the video and records its position in each frame, allowing it to track the child's position even if the child moves within the video.

[0538] 5. Automatic generation of video clips

[0539] The server automatically extracts the footage containing the subject and generates individual video clips using video editing software such as FFmpeg, and automatically applies zoom and frame correction to create visually pleasing clips.

[0540] 6. Generating video files

[0541] The server then assembles the edited video clips into a series of video files, a process that organizes the clips chronologically to create a coherent video.

[0542] 7. Upload to cloud storage

[0543] The server then uploads the generated video files to cloud storage, typically a cloud-based storage service such as Google Drive or Dropbox.

[0544] 8. Providing a User Interface

[0545] Users access video files stored in cloud storage through a dedicated application, which is intuitive and easy to use, designed to allow users to easily find, download, and watch videos.

[0546] Specific examples

[0547] In the case of sports day

[0548] 1. Event setup: The user sets up the "Sports Day" event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0549] 2. Start shooting: The user sets up a high-resolution camera at the sports day venue and starts shooting from a position that covers the entire venue.

[0550] 3. Real-time video transmission: The device transmits video data to the server in real time. The server analyzes the received video and performs facial recognition and tracking of the subject.

[0551] 4. Clip generation and video editing: The server automatically extracts and edits the video clips containing the target children. For example, it clips the footage of the children running and edits it to include all the necessary scenes.

[0552] 5. Video file distribution: The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0553] Prompt Sentence Examples

[0554] markdown

[0555] Describe a system that automatically edits and saves footage of children's sports events to cloud storage. The user installs a high-resolution camera and configures the event using a dedicated app. The device transmits the footage in real time to a server, which uses AI facial recognition to identify the children and automatically edits the footage. The final video file is uploaded to cloud storage and can be accessed by the user.

[0556] This system allows parents to record and enjoy high-quality footage of their children's growth and memories without the effort of filming and editing during events. By increasing the amount of time parents spend cheering and interacting with their children at the venue, parents can spend more fulfilling time with their children.

[0557] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0558] Step 1:

[0559] The user launches a dedicated application and sets up the event. First, they move to the settings screen within the application and enter the event name (e.g., "Sports Day") and date. Next, they upload a facial image of the child in question. This facial image will be used for facial recognition later. The event information and facial image set by the user are entered. The entered data is saved within the application.

[0560] Step 2:

[0561] Users set up high-resolution cameras at the event venue. They position the cameras so that they can cover the entire event, and adjust the angle and height as needed. When they're ready to start recording, they press the camera's record button to begin recording. The cameras generate high-quality video data in real time, allowing the entire event to be recorded in detail.

[0562] Step 3:

[0563] The terminal (device connected to the camera) compresses the captured video data in real time and sends it to a server via the Internet. Video data is input and output in a compressed form. The video data sent from the terminal is received by the server. This process achieves efficient data transfer.

[0564] Step 4:

[0565] The server analyzes the received video data and uses an AI facial recognition algorithm (such as Amazon Rekognition or OpenCV) to match it with previously uploaded facial images. The facial image is input and the face of a specific child is recognized. The server identifies the target person based on this processing. The location of the specific child within the video is confirmed through facial recognition.

[0566] Step 5:

[0567] The server tracks the recognized subject's face in the video and records their location in each frame. The location information obtained through facial recognition is input, and tracking data including the subject's location is output. The server can use this tracking data to track the location of a specific child over time.

[0568] Step 6:

[0569] The server automatically extracts the portion of the video that shows the subject and generates a separate video clip. This process is performed using video editing software (e.g., FFmpeg). Tracking data is input and a video clip with the subject at the center is output. The server then applies zoom and frame correction to this clip to make it more visually appealing.

[0570] Step 7:

[0571] The server chronologically organizes multiple video clips into a series of video files. Edited clips are input and a coherent video file is output. The server encodes this video file and converts it to the optimal format (e.g., MP4).

[0572] Step 8:

[0573] The server uploads the generated video file to the cloud storage specified by the user (e.g., Google Drive or Dropbox). The completed video file is input and saved in the cloud storage. The upload process sends the file over the Internet and stores it securely in the cloud.

[0574] Step 9:

[0575] Users access cloud storage through a dedicated application to view uploaded video files. Users can search for video files and download or stream them. The application provides an intuitive interface, making it easy to find the video you want.

[0576] (Application example 1)

[0577] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0578] In conventional virtual shopping, parents must manually record and then edit videos to record their children's real-time reactions. This process requires time and effort for parents, hindering an intuitive and natural shopping experience. Furthermore, manual filming and editing increases the risk of missing important moments. Therefore, there is a need for a system that can automatically capture and record children's reactions and fun moments.

[0579] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0580] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, and means for automatically capturing and recording a child's reactions and moments of enjoyment during virtual shopping. This eliminates the need for parents to manually shoot and edit videos, and allows parents to easily record important moments during virtual shopping without missing them.

[0581] A "high resolution image capture device" is a camera device capable of capturing video at high resolution.

[0582] "Means for transmitting to the server in real time" refers to technology that transfers video data to the server immediately without delay.

[0583] "Means for identifying a subject using a facial recognition algorithm on a server" refers to a method in which a server uses facial recognition technology to identify a specific person.

[0584] "Means for tracking a subject in a video" refers to a technology that continuously tracks the position and movements of a specific person in a video.

[0585] "Means for automatically generating clips" refers to a technology that automatically cuts out specific parts of video and edits them into short video segments.

[0586] The "means for generating a video file" refers to a method for integrating multiple video clips and editing them into a single continuous video file.

[0587] "Means for uploading to cloud storage" refers to a technology that stores the generated video files on a remote server and makes them accessible via the Internet.

[0588] "Means for providing a user interface" refers to operation screens and applications that allow users to interact with the system intuitively.

[0589] "Virtual shopping" is the act of selecting and purchasing products in a virtual space using the Internet.

[0590] "A means for automatically capturing and recording children's reactions and moments of enjoyment" is a technology that detects children's facial expressions and behavior in real time and automatically records those moments as video.

[0591] This invention is a system that automatically captures and records moments that parents enjoy with their children during virtual shopping, allowing parents to preserve memories without any hassle.

[0592] System configuration

[0593] 1. High-resolution imaging equipment

[0594] Users place a high-resolution camera in the virtual shopping environment to capture detailed, high-resolution footage of the entire shopping experience. The camera remains stationary and captures footage of the shopping experience.

[0595] 2. Real-time video transmission method

[0596] The device (smartphone or head-mounted display) transmits the captured video data to a server in real time using an internet connection, and video data captured while shopping is sent to the server instantly.

[0597] 3. Facial Recognition and Video Analysis

[0598] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0599] 4. How to track the subject

[0600] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0601] 5. Automatic Clip Generation Method

[0602] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0603] 6. Video file generation method

[0604] The server compiles the video clips into a series of video files, which are organized chronologically to create a coherent video of the entire shopping experience.

[0605] 7. Cloud Storage Distribution Methods

[0606] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0607] 8. User Interface

[0608] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0609] Hardware and software used

[0610] Hardware

[0611] Webcam

[0612] Smartphone

[0613] head-mounted display

[0614] software

[0615] OpenCV (image processing library)

[0616] dlib (face recognition algorithm)

[0617] moviepy (video editing library)

[0618] Specific application examples

[0619] The case for virtual shopping

[0620] 1. Pre-registration of facial images

[0621] Users upload a photo of their child's face in advance through a dedicated application.

[0622] 2. Start shooting

[0623] Users place high-resolution cameras in the virtual shopping environment and capture footage in real time.

[0624] 3. Video data transmission

[0625] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0626] 4. Clip Creation and Video Editing

[0627] The server automatically cuts out clips of footage in which the child appears, and edits them together, for example, when the child shows interest in a product, and edits the clips to include all the important moments.

[0628] 5. Distribution of video files

[0629] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0630] Prompt Sentence Examples

[0631] "Develop an application that automatically captures a child's smile or surprised expression while they are shopping, and creates a video of their memories that can be viewed later."

[0632] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0633] Step 1:

[0634] Users first upload a facial image of their child through a dedicated application. In this step, users register their child's facial image data in the system using a smartphone or computer. The system receives this image data as input and generates and stores facial feature data for use in the facial recognition algorithm.

[0635] Step 2:

[0636] The user installs a high-resolution camera in the virtual shopping environment and captures video in real time. The video captured by the camera is positioned to cover the entire view of the shop. The camera captures video data in real time and transmits it to the terminal. The terminal then transmits the received video data to the server in real time.

[0637] Step 3:

[0638] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face. In this step, the server receives real-time video data as input and runs the facial recognition algorithm. The algorithm compares the face in the real-time video with pre-registered facial feature data, and if there is a match, the child's face is identified.

[0639] Step 4:

[0640] The server tracks the recognized subject's face in the video and records its position in each frame. In this step, the server continues to record the position of the identified face in each frame of the video. For each frame received as input, it provides the face position information as output.

[0641] Step 5:

[0642] The server automatically extracts the video portion where the subject appears and generates a separate video clip. In this step, the server analyzes the video based on the facial position information, extracts the portion where the subject appears, and creates a clip. Zoom and frame correction are also applied automatically. The input is the tracked video frame and position information, and the output is a visually pleasing video clip.

[0643] Step 6:

[0644] The server edits multiple video clips and compiles them into a series of video files. In this step, the server integrates the generated video clips to create a time-coherent video file. It takes individual video clips as input, edits them, and generates a continuous video file as output.

[0645] Step 7:

[0646] The server uploads the generated video file to the cloud storage. In this step, the server saves the completed video file to the cloud storage via the Internet. It receives the completed video file as input and provides the cloud storage URL or access information as output.

[0647] Step 8:

[0648] Users access, download, and watch video files stored in cloud storage through a dedicated application. In this step, users access cloud storage using a dedicated user interface to view and save video files. Cloud storage access information is used as input, and video files are made available for viewing as output.

[0649] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0650] The present invention is a system that automatically generates individual videos for families with children and then uses an emotion engine to edit the videos based on the user's emotions. This system eliminates the need for parents to film and edit events, allowing them to easily record their children's memories and highlight moving moments. Specific embodiments of the system are described below.

[0651] System configuration

[0652] 1. High-resolution imaging equipment

[0653] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[0654] 2. Real-time video transmission method

[0655] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[0656] 3. Facial Recognition and Video Analysis

[0657] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0658] 4. How to track the subject

[0659] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0660] 5. Automatic Clip Generation Method

[0661] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0662] 6. Video file generation method

[0663] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[0664] 7. Cloud Storage Distribution Methods

[0665] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0666] 8. User Interface

[0667] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0668] Incorporating an emotion engine

[0669] 1. Emotion Engine

[0670] The server also has an emotion engine that recognizes the user's emotions from the video, analyzing emotions such as smile, sadness, and surprise in real time to obtain emotional data during the event.

[0671] 2. Emotional editing

[0672] The server then edits the video clips based on the recognized emotions, for example, prioritizing scenes with many smiling faces as highlight clips, or editing to emphasize moving scenes.

[0673] Specific examples

[0674] Emotion Recognition in the Case of Sports Day

[0675] 1. Event Settings

[0676] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0677] 2. Start shooting

[0678] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0679] 3. Video data transmission

[0680] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0681] 4. Emotion analysis

[0682] The server also uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[0683] 5. Clip Creation and Video Editing

[0684] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[0685] 6. Distribution of video files

[0686] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0687] This system allows users to record their children's growth and memories in a richer way, without the hassle of filming and editing during events, and enjoy footage that highlights moving moments. This increases the amount of time parents and children can spend together cheering and interacting with each other at the venue, allowing for more fulfilling parent-child time.

[0688] The processing flow will be explained below.

[0689] Step 1:

[0690] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[0691] Step 2:

[0692] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[0693] Step 3:

[0694] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[0695] Step 4:

[0696] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[0697] Step 5:

[0698] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0699] Step 6:

[0700] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[0701] Step 7:

[0702] The server also uses an emotion engine to analyze the emotions of the subjects in the video and the viewers in real time, detecting, for example, the smile or other emotions of a child crossing the finish line in a race.

[0703] Step 8:

[0704] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[0705] Step 9:

[0706] The server determines the importance of each clip based on the emotional data recognized by the emotion engine, and prioritizes highlights, such as scenes of smiling faces or cheering.

[0707] Step 10:

[0708] The server then edits the selected video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[0709] Step 11:

[0710] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[0711] Step 12:

[0712] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[0713] This series of processes eliminates the need for users to perform tedious filming and editing tasks during events, allowing them to spend more time enjoying their children's growth and memories. Furthermore, the emotion engine makes it possible to create videos that emphasize moving moments.

[0714] Example 2

[0715] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] For families with children, the time and effort required to film and edit events is a challenge. In particular, if parents are too focused on filming during an event, they lose the time they need to cheer on and interact with their children in real time. Furthermore, editing the footage to appropriately highlight moving moments and records of children's growth is difficult.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0718] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject in the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for recognizing emotions in the video using an emotion engine, and means for editing the video based on the recognized emotions. This allows users to easily create videos that highlight their children's growth and moving moments without the hassle of filming and editing the event.

[0719] "High-resolution image capture device" refers to a capture device capable of capturing detailed footage of the entire event.

[0720] "Means for transmitting to a server in real time" refers to a communication means for immediately transferring captured video data to a server.

[0721] A "facial recognition algorithm" is software or a program that identifies the faces of people in a video and authenticates a specific individual.

[0722] "Subject" refers to an individual identified by a facial recognition algorithm.

[0723] A "tracking means" is a system for continuously recording the location of a subject recognized in the video.

[0724] The "means for automatically generating clips" is a function that automatically cuts out the portion of the video in which the subject appears.

[0725] The "means for generating a video file" is a system that combines multiple video clips and edits them into a single video file.

[0726] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.

[0727] "User interface" refers to the screen and input means used by users to operate software or a system.

[0728] An "emotion engine" is software or an algorithm that analyzes the emotions of people in a video and recognizes their emotional state.

[0729] "Means for editing video" refers to a function for processing captured video data and emphasizing or changing content based on specific scenes or emotions.

[0730] This system automatically generates individual videos for families with children and uses an emotion engine to edit the videos based on the user's emotions. This saves parents the trouble of filming and editing events, allowing them to easily record their children's memories and highlight moving moments.

[0731] This system is implemented using the following hardware and software.

[0732] High-resolution image capture device: The user installs a high-resolution camera at the event venue to capture the entire event. For example, the Sony Alpha series is used.

[0733] Real-time video transmission method: The device (e.g., laptop or Raspberry Pi) transmits the video data received from the camera to the server in real time. The communication method is a high-speed Internet connection (Wi-Fi or wired LAN).

[0734] Facial recognition algorithm: The server uses an AI facial recognition algorithm such as Amazon Rekognition to recognize the child's face from the received video data and identify the child by comparing it with previously uploaded facial image data.

[0735] Subject tracking method: The server continuously tracks the location information of the recognized subject within the video. It records the position in each frame, so it can follow the subject even if they move within the video.

[0736] Automatic clip generation: The server automatically extracts the video portion in which the recognized subject appears and generates individual video clips. The clips are then edited to be visually appealing, with automatic zooming and frame correction.

[0737] Video file generator: The server edits multiple video clips and compiles them into a series of video files, creating a consistent video of the entire event.

[0738] Cloud storage distribution method: The server uploads the completed video file to cloud storage (e.g., Google Drive). The data is saved in the cloud storage specified by the user.

[0739] User interface: Users can access cloud storage through a dedicated app (iOS or Android app) and download or watch video files.

[0740] Emotion engine: The server uses Amazon Rekognition and Microsoft Azure Emotion API to analyze the user's emotions in real time. It recognizes emotions such as smiles and surprise and records data during the event.

[0741] Emotion-based editing: The server edits video clips based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights and emphasizes moving scenes.

[0742] Example: Emotion recognition in the case of athletic meet

[0743] 1. Event Settings

[0744] Users set up a "Sports Day" event using a dedicated app, and upload facial images of the children involved in the event in advance.

[0745] 2. Start shooting

[0746] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0747] 3. Video data transmission

[0748] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0749] 4. Emotion analysis

[0750] The server uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[0751] 5. Clip Creation and Video Editing

[0752] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[0753] 6. Distribution of video files

[0754] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated app.

[0755] Example prompts for generative AI models:

[0756] Perform emotion analysis on video data from a sports day and generate a video clip that emphasizes touching moments. The facial images of the target children and the actual video data can be found at the following links: (Facial image URL), (Video data URL). Edit the scenes in which smiling faces are detected as highlights.

[0757] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0758] Step 1:

[0759] The user sets up a high-resolution camera at the event venue, fixes the camera in a position where it can capture the entire event in detail, and starts recording, inputting video data of the entire event.

[0760] Step 2:

[0761] The device (laptop or Raspberry Pi) transmits the video data acquired from the camera to the server in real time. The communication method is a high-speed internet connection. The input in this step is the video data from the camera, and the output is the real-time video data transmission to the server.

[0762] Step 3:

[0763] The server analyzes the received video data using AI facial recognition algorithms such as Amazon Rekognition. It compares the data with previously uploaded facial image data and recognizes the child's face. The input for this step is real-time video data and facial image data, and the output is the location information of the recognized face.

[0764] Step 4:

[0765] The server tracks the recognized face in the video and keeps recording the subject's position information for each frame. The input is the position information of the recognized face, and the output is the subject's position data for each frame.

[0766] Step 5:

[0767] The server automatically extracts the portion of the video in which the recognized and tracked subject appears, generating individual video clips. The clips are then edited for visual clarity with automatic zoom and frame correction. The inputs are the subject's position data and video data, and the output is the video clip.

[0768] Step 6:

[0769] The server then edits the generated video clips into a series of video files. The clips are organized in chronological order to create a seamless, coherent video. The input is a series of video clips, and the output is an edited video file.

[0770] Step 7:

[0771] The server uploads the completed video file to the specified cloud storage, such as Google Drive or Amazon S3. The input is the edited video file, and the output is a notification that the upload to the cloud storage is complete.

[0772] Step 8:

[0773] Users use a dedicated app to access cloud storage and download or view video files. The interface is intuitive and easy to use. The input is the cloud storage URL, and the output is a video file that can be downloaded and viewed.

[0774] Step 9:

[0775] The server uses an emotion engine (such as Amazon Rekognition or Microsoft Azure Emotion API) to analyze the user's emotions in the video in real time. It recognizes emotions such as smiles and surprise and records emotional data during the event. The input is the video data, and the output is the emotion recognition results.

[0776] Step 10:

[0777] The server then edits the video clip based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights, and further emphasizes moving scenes. The input is the emotion recognition results and the video clip, and the output is an edited video based on the emotions.

[0778] (Application example 2)

[0779] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0780] In modern factory environments, worker safety management is a critical issue. Conventional monitoring methods have difficulty monitoring in real time whether workers are working in an appropriate state of health or under emotional stress. Furthermore, dangerous situations frequently occur at worksites due to the large number of machines in operation. Conventional safety monitoring systems rely on simple surveillance cameras, which are unable to analyze workers' emotions or fatigue levels. This makes it difficult to adequately ensure safety in the actual work environment. Therefore, there is a need for a system that can analyze workers' emotional state and fatigue levels in real time and respond quickly.

[0781] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video using a high-resolution image capture device, means for transmitting the video in real time, means for identifying a target person using a facial recognition algorithm, means for tracking the target person in the video, means for automatically generating clips, means for editing the clips to generate a video file, means for uploading the video file to cloud storage, means for providing a user interface for accessing the video file, means for analyzing emotions in the video in real time using an emotion engine and issuing a safety warning, and means for transmitting the analysis data in real time, and means for storing the data in cloud storage. This makes it possible to monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[0782] A "high resolution imaging device" is an imaging device for capturing highly detailed videos and images.

[0783] "Means for transmitting to the server in real time" refers to the communication technology for instantly sending captured video data to the server.

[0784] A "face recognition algorithm" is a computer program that identifies faces in video and compares them with pre-registered facial information.

[0785] "Means for tracking a subject within a video" refers to a technology that tracks the position of a specific person even if they move within the video.

[0786] "Means for automatically generating clips" refers to technology that automatically cuts out specific video segments and generates them as short video clips.

[0787] The "means for generating a video file" refers to the technology for editing clips and putting them together as a series of videos.

[0788] "Means for uploading to cloud storage" refers to a technology for storing the generated video file in a remote storage area on the Internet.

[0789] A "user interface" is the portion of software that provides the visual and operational elements through which a user interacts with a system.

[0790] An "emotion engine" is an algorithm and software that analyzes the emotions of people in a video and identifies their emotional state.

[0791] "Means for issuing safety warnings" refers to technology for issuing warnings when dangerous situations are detected as a result of emotion analysis.

[0792] "Data transmission means" refers to a communication technique for sending the analyzed data to another location.

[0793] "Means by which data is stored in cloud storage" refers to the technology used to store analyzed data and generated videos in the cloud.

[0794] The present invention relates to a worker safety monitoring system for a factory, which uses a high-resolution image capture device, a high-performance server, and cloud storage. The system analyzes the facial expressions and postures of workers working in the factory and can perform safety management in real time. Specific embodiments of the system are described below.

[0795] System configuration

[0796] 1. High-resolution imaging equipment

[0797] The user installs a high-resolution camera in the factory, which covers a wide area and captures detailed images of workers. For example, the Logitech C920 is used as this camera.

[0798] 2. Real-time video transmission method

[0799] The device transmits the captured video data in real time to a server using an internet connection, where the video data is processed immediately.

[0800] 3. Facial Recognition Algorithm

[0801] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the worker's face. At this time, it compares it with pre-registered facial images to identify the specific worker. OpenCV is used for facial recognition.

[0802] 4. How to track the subject

[0803] The server tracks the recognized worker's face in the video and records the position information in each frame, making it possible to track the target's position even if the target moves within the video.

[0804] 5. Emotion analysis

[0805] The server uses an emotion engine to analyze the emotions of workers in the video in real time. For example, it analyzes emotions such as smiles, sadness, and surprise, and generates data useful for safety management. An AI model using TensorFlow / Keras is used for emotion analysis.

[0806] 6. Issuance of safety warnings

[0807] Based on the analysis results of the emotion engine, the server quickly issues safety warnings, including visual alerts and audio notifications, if a worker is in a dangerous or overly fatigued state.

[0808] 7. Data transmission and storage in cloud storage

[0809] The server uploads the analyzed data and generated video clips to cloud storage and also transmits the analyzed data to other devices in real time, allowing factory managers to monitor the safety status of workers even from remote locations.

[0810] 8. User Interface

[0811] Users can access cloud storage and check saved data and video clips using a dedicated application. The intuitive user interface allows users to quickly obtain the information they need.

[0812] Specific examples

[0813] To implement a system to monitor the safety of workers in a factory, high-resolution cameras are installed in each work area. The cameras capture the facial expressions and postures of workers in real time and send the footage to a server. The server then uses a facial recognition algorithm and an emotion engine to identify the worker and perform emotion analysis. Based on the analysis results, a safety warning is issued immediately if the worker is in danger. Furthermore, the real-time analysis data is stored in cloud storage and can be accessed by managers through a dedicated application.

[0814] Example prompts for generative AI models

[0815] Design a system that uses facial recognition and emotion analysis to monitor the safety of workers in a factory. A high-resolution camera captures images of workers working and analyzes them in real time using an emotion engine. If a worker is in danger or overly fatigued, an alert will be issued and the data will be stored in the cloud. Please provide a concrete code example.

[0816] As a result, a system can be provided that can monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[0817] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0818] Step 1:

[0819] Users install high-resolution cameras in the work area of ​​their factories, which capture detailed images of workers' faces and the work they are doing.

[0820] Input: spatial information (camera installation position), high-resolution camera

[0821] Output: Real-time video data

[0822] Step 2:

[0823] The device transmits the captured real-time video data to a server via the Internet, where streaming technology is used to efficiently transmit large amounts of data.

[0824] Input: Real-time video data

[0825] Output: Video data sent to the server

[0826] Step 3:

[0827] The server uses OpenCV to analyze the received video data and identifies the worker's face using a facial recognition algorithm, which then compares it with pre-registered facial information to identify each worker.

[0828] Input: Transmitted video data, pre-registered face data

[0829] Output: Video data including identified workers (with location information)

[0830] Step 4:

[0831] The server tracks the face of the identified worker and records their position in each frame, allowing the worker to be tracked in real time even if they move within the camera's field of view.

[0832] Input: Video data containing identified workers

[0833] Output: Video data with location information of tracked workers

[0834] Step 5:

[0835] The server automatically generates a clip of the worker's face, zooming and adjusting the frame as needed, using an AI model to generate the optimal clip.

[0836] Input: Location-based video data of tracked workers

[0837] Output: Automatically generated clip video data

[0838] Step 6:

[0839] The server uses an emotion engine to analyze the emotions of the workers in the video clips, specifically identifying emotional states such as smiling or surprised using TensorFlow / Keras.

[0840] Input: Clip video data

[0841] Output: Parsed emotion data (with emotion labels)

[0842] Step 7:

[0843] The server then issues safety alerts based on the analyzed emotion data, for example, visual and audio alerts if a worker is in a dangerous emotional state (anger or sadness).

[0844] Input: Emotion-labeled analysis data

[0845] Output: Safety warning (alert)

[0846] Step 8:

[0847] The server uploads the analyzed data and generated video clips to cloud storage, and also transmits the analyzed data in real time to a remote administrator terminal.

[0848] Input: Analysis data, clip video data

[0849] Output: Data stored in the cloud, analysis data sent to the administrator's terminal

[0850] Step 9:

[0851] Users (administrators) can access cloud storage and check stored data and video clips using a dedicated application. The application has an intuitive interface, allowing users to quickly obtain the information they need.

[0852] Input: Data stored in the cloud (video clips, analysis data)

[0853] Output: Monitoring information available to administrators (interface operations)

[0854] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0855] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0856] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0857] [Third embodiment]

[0858] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0859] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0860] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0861] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0862] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0863] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0864] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0865] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0866] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0867] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0868] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0869] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0870] The present invention is a system for automatically generating personalized videos for families with children, allowing parents to easily record their children's memories without the need to film and edit events. This system is implemented as follows.

[0871] System configuration

[0872] 1. High-resolution imaging equipment

[0873] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[0874] 2. Real-time video transmission method

[0875] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[0876] 3. Facial Recognition and Video Analysis

[0877] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0878] 4. How to track the subject

[0879] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[0880] 5. Automatic Clip Generation Method

[0881] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[0882] 6. Video file generation method

[0883] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[0884] 7. Cloud Storage Distribution Methods

[0885] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[0886] 8. User Interface

[0887] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[0888] Specific examples

[0889] Sports day case

[0890] 1. Event Settings

[0891] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0892] 2. Start shooting

[0893] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[0894] 3. Video data transmission

[0895] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[0896] 4. Clip Creation and Video Editing

[0897] The server automatically extracts and edits the footage of the child in question into a clip. For example, a clip of the child running is clipped and edited to include all the necessary scenes.

[0898] 5. Distribution of video files

[0899] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0900] This system allows parents to easily record and enjoy their children's growth and memories without the hassle of filming and editing during events, increasing the amount of time parents can spend cheering and interacting with their children at the venue and creating a more fulfilling parent-child experience.

[0901] The processing flow will be explained below.

[0902] Step 1:

[0903] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[0904] Step 2:

[0905] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[0906] Step 3:

[0907] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[0908] Step 4:

[0909] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[0910] Step 5:

[0911] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[0912] Step 6:

[0913] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[0914] Step 7:

[0915] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[0916] Step 8:

[0917] The server then edits the multiple video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[0918] Step 9:

[0919] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[0920] Step 10:

[0921] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[0922] This series of processes eliminates the need for the user to perform tedious shooting and editing tasks during the event, allowing them to have more time to enjoy their children's growth and memories.

[0923] Example 1

[0924] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0925] Nowadays, parents spend a lot of time and effort recording their children's growth and important events. Filming and editing events is particularly tedious work, requiring parents to spend a lot of time on the task, which can result in parents being unable to fully enjoy the event itself. Furthermore, the quality of the footage recorded varies depending on the cameraman's skill and the equipment used. Therefore, there is a need for a system that allows parents to easily obtain high-quality footage.

[0926] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0927] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video in real time to a data processing device, means for identifying a subject using a facial recognition algorithm on the data processing device, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for automatically zooming and frame correction for the subject, and means for pre-registering facial images of the subject. This allows parents to easily record and enjoy high-quality video of their children's growth and activities during the event without the effort of filming or editing.

[0928] "High-resolution image capturing device" refers to any imaging device that has a high pixel count and can capture detailed, clear images.

[0929] "Means for transmitting video in real time" refers to communication technology and devices for transferring captured video data to a server or data processing device immediately without delay.

[0930] A "facial recognition algorithm" refers to software or computational methods used to identify faces in images and detect specific individuals.

[0931] "Means for identifying subjects" refers to functions and technologies for recognizing and identifying individual people within video footage.

[0932] "Means for tracking a subject within a video stream" refers to the technology or computational methods used to continuously track the location of an identified person as they move within the video stream.

[0933] "Means for automatically generating clips" refers to functions and technologies for cutting out footage showing specific people and automatically editing it into short video segments.

[0934] "Means for generating individual video files" refers to the functions and technologies for editing and integrating multiple video clips into a continuous video file.

[0935] "Means for uploading to cloud storage" refers to the functionality and technology for storing the generated video files in a cloud-based storage device via the Internet.

[0936] "Means for providing a user interface" refers to display screens and operating means that make it easier for users to operate systems and services.

[0937] "Means for automatic zoom and frame correction" refers to functions and technologies for zooming in on a specific subject in a video or automatically adjusting the composition of the video.

[0938] "Means for pre-registering facial images" refers to functions and technologies that allow the system to load and store facial image data of the subject in advance.

[0939] MODE FOR CARRYING OUT THE INVENTION

[0940] This invention is an automatic individual video generation system designed for families with children, which reduces the effort required for parents to film and edit events and allows them to easily record their children's growth and memories. This system is implemented using the following configuration and process.

[0941] System configuration

[0942] 1. High-resolution imaging equipment

[0943] Users install a high-resolution camera at the event venue. This camera takes still pictures and captures the entire event in detail. This device can be, for example, a 4K video camera or a high-performance digital SLR camera.

[0944] 2. Real-time video transmission

[0945] The device transmits the captured video data to a server in real time. This process uses an internet connection, so the data is sent to the server instantly. Specifically, a smartphone or tablet receives the data from the camera and sends it to the server via Wi-Fi or mobile network.

[0946] 3. Facial Recognition and Video Analysis

[0947] The server analyzes the received video data and uses AI facial recognition algorithms to identify the specific child's face, using facial recognition software such as Amazon Rekognition or OpenCV.

[0948] 4. Subject Tracking

[0949] The server tracks the recognized child's face in the video and records its position in each frame, allowing it to track the child's position even if the child moves within the video.

[0950] 5. Automatic generation of video clips

[0951] The server automatically extracts the footage containing the subject and generates individual video clips using video editing software such as FFmpeg, and automatically applies zoom and frame correction to create visually pleasing clips.

[0952] 6. Generating video files

[0953] The server then assembles the edited video clips into a series of video files, a process that organizes the clips chronologically to create a coherent video.

[0954] 7. Upload to cloud storage

[0955] The server then uploads the generated video files to cloud storage, typically a cloud-based storage service such as Google Drive or Dropbox.

[0956] 8. Providing a User Interface

[0957] Users access video files stored in cloud storage through a dedicated application, which is intuitive and easy to use, designed to allow users to easily find, download, and watch videos.

[0958] Specific examples

[0959] In the case of sports day

[0960] 1. Event setup: The user sets up the "Sports Day" event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[0961] 2. Start shooting: The user sets up a high-resolution camera at the sports day venue and starts shooting from a position that covers the entire venue.

[0962] 3. Real-time video transmission: The device transmits video data to the server in real time. The server analyzes the received video and performs facial recognition and tracking of the subject.

[0963] 4. Clip generation and video editing: The server automatically extracts and edits the video clips containing the target children. For example, it clips the footage of the children running and edits it to include all the necessary scenes.

[0964] 5. Video file distribution: The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[0965] Prompt Sentence Examples

[0966] markdown

[0967] Describe a system that automatically edits and saves footage of children's sports events to cloud storage. The user installs a high-resolution camera and configures the event using a dedicated app. The device transmits the footage in real time to a server, which uses AI facial recognition to identify the children and automatically edits the footage. The final video file is uploaded to cloud storage and can be accessed by the user.

[0968] This system allows parents to record and enjoy high-quality footage of their children's growth and memories without the effort of filming and editing during events. By increasing the amount of time parents spend cheering and interacting with their children at the venue, parents can spend more fulfilling time with their children.

[0969] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0970] Step 1:

[0971] The user launches a dedicated application and sets up the event. First, they move to the settings screen within the application and enter the event name (e.g., "Sports Day") and date. Next, they upload a facial image of the child in question. This facial image will be used for facial recognition later. The event information and facial image set by the user are entered. The entered data is saved within the application.

[0972] Step 2:

[0973] Users set up high-resolution cameras at the event venue. They position the cameras so that they can cover the entire event, and adjust the angle and height as needed. When they're ready to start recording, they press the camera's record button to begin recording. The cameras generate high-quality video data in real time, allowing the entire event to be recorded in detail.

[0974] Step 3:

[0975] The terminal (device connected to the camera) compresses the captured video data in real time and sends it to a server via the Internet. Video data is input and output in a compressed form. The video data sent from the terminal is received by the server. This process achieves efficient data transfer.

[0976] Step 4:

[0977] The server analyzes the received video data and uses an AI facial recognition algorithm (such as Amazon Rekognition or OpenCV) to match it with previously uploaded facial images. The facial image is input and the face of a specific child is recognized. The server identifies the target person based on this processing. The location of the specific child within the video is confirmed through facial recognition.

[0978] Step 5:

[0979] The server tracks the recognized subject's face in the video and records their location in each frame. The location information obtained through facial recognition is input, and tracking data including the subject's location is output. The server can use this tracking data to track the location of a specific child over time.

[0980] Step 6:

[0981] The server automatically extracts the portion of the video that shows the subject and generates a separate video clip. This process is performed using video editing software (e.g., FFmpeg). Tracking data is input and a video clip with the subject at the center is output. The server then applies zoom and frame correction to this clip to make it more visually appealing.

[0982] Step 7:

[0983] The server chronologically organizes multiple video clips into a series of video files. Edited clips are input and a coherent video file is output. The server encodes this video file and converts it to the optimal format (e.g., MP4).

[0984] Step 8:

[0985] The server uploads the generated video file to the cloud storage specified by the user (e.g., Google Drive or Dropbox). The completed video file is input and saved in the cloud storage. The upload process sends the file over the Internet and stores it securely in the cloud.

[0986] Step 9:

[0987] Users access cloud storage through a dedicated application to view uploaded video files. Users can search for video files and download or stream them. The application provides an intuitive interface, making it easy to find the video you want.

[0988] (Application example 1)

[0989] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0990] In conventional virtual shopping, parents must manually record and then edit videos to record their children's real-time reactions. This process requires time and effort for parents, hindering an intuitive and natural shopping experience. Furthermore, manual filming and editing increases the risk of missing important moments. Therefore, there is a need for a system that can automatically capture and record children's reactions and fun moments.

[0991] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0992] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, and means for automatically capturing and recording a child's reactions and moments of enjoyment during virtual shopping. This eliminates the need for parents to manually shoot and edit videos, and allows parents to easily record important moments during virtual shopping without missing them.

[0993] A "high resolution image capture device" is a camera device capable of capturing video at high resolution.

[0994] "Means for transmitting to the server in real time" refers to technology that transfers video data to the server immediately without delay.

[0995] "Means for identifying a subject using a facial recognition algorithm on a server" refers to a method in which a server uses facial recognition technology to identify a specific person.

[0996] "Means for tracking a subject in a video" refers to a technology that continuously tracks the position and movements of a specific person in a video.

[0997] "Means for automatically generating clips" refers to a technology that automatically cuts out specific parts of video and edits them into short video segments.

[0998] The "means for generating a video file" refers to a method for integrating multiple video clips and editing them into a single continuous video file.

[0999] "Means for uploading to cloud storage" refers to a technology that stores the generated video files on a remote server and makes them accessible via the Internet.

[1000] "Means for providing a user interface" refers to operation screens and applications that allow users to interact with the system intuitively.

[1001] "Virtual shopping" is the act of selecting and purchasing products in a virtual space using the Internet.

[1002] "A means for automatically capturing and recording children's reactions and moments of enjoyment" is a technology that detects children's facial expressions and behavior in real time and automatically records those moments as video.

[1003] This invention is a system that automatically captures and records moments that parents enjoy with their children during virtual shopping, allowing parents to preserve memories without any hassle.

[1004] System configuration

[1005] 1. High-resolution imaging equipment

[1006] Users place a high-resolution camera in the virtual shopping environment to capture detailed, high-resolution footage of the entire shopping experience. The camera remains stationary and captures footage of the shopping experience.

[1007] 2. Real-time video transmission method

[1008] The device (smartphone or head-mounted display) transmits the captured video data to a server in real time using an internet connection, and video data captured while shopping is sent to the server instantly.

[1009] 3. Facial Recognition and Video Analysis

[1010] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1011] 4. How to track the subject

[1012] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[1013] 5. Automatic Clip Generation Method

[1014] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[1015] 6. Video file generation method

[1016] The server compiles the video clips into a series of video files, which are organized chronologically to create a coherent video of the entire shopping experience.

[1017] 7. Cloud Storage Distribution Methods

[1018] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[1019] 8. User Interface

[1020] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[1021] Hardware and software used

[1022] Hardware

[1023] Webcam

[1024] Smartphone

[1025] head-mounted display

[1026] software

[1027] OpenCV (image processing library)

[1028] dlib (face recognition algorithm)

[1029] moviepy (video editing library)

[1030] Specific application examples

[1031] The case for virtual shopping

[1032] 1. Pre-registration of facial images

[1033] Users upload a photo of their child's face in advance through a dedicated application.

[1034] 2. Start shooting

[1035] Users place high-resolution cameras in the virtual shopping environment and capture footage in real time.

[1036] 3. Video data transmission

[1037] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1038] 4. Clip Creation and Video Editing

[1039] The server automatically cuts out clips of footage in which the child appears, and edits them together, for example, when the child shows interest in a product, and edits the clips to include all the important moments.

[1040] 5. Distribution of video files

[1041] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1042] Prompt Sentence Examples

[1043] "Develop an application that automatically captures a child's smile or surprised expression while they are shopping, and creates a video of their memories that can be viewed later."

[1044] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1045] Step 1:

[1046] Users first upload a facial image of their child through a dedicated application. In this step, users register their child's facial image data in the system using a smartphone or computer. The system receives this image data as input and generates and stores facial feature data for use in the facial recognition algorithm.

[1047] Step 2:

[1048] The user installs a high-resolution camera in the virtual shopping environment and captures video in real time. The video captured by the camera is positioned to cover the entire view of the shop. The camera captures video data in real time and transmits it to the terminal. The terminal then transmits the received video data to the server in real time.

[1049] Step 3:

[1050] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face. In this step, the server receives real-time video data as input and runs the facial recognition algorithm. The algorithm compares the face in the real-time video with pre-registered facial feature data, and if there is a match, the child's face is identified.

[1051] Step 4:

[1052] The server tracks the recognized subject's face in the video and records its position in each frame. In this step, the server continues to record the position of the identified face in each frame of the video. For each frame received as input, it provides the face position information as output.

[1053] Step 5:

[1054] The server automatically extracts the video portion where the subject appears and generates a separate video clip. In this step, the server analyzes the video based on the facial position information, extracts the portion where the subject appears, and creates a clip. Zoom and frame correction are also applied automatically. The input is the tracked video frame and position information, and the output is a visually pleasing video clip.

[1055] Step 6:

[1056] The server edits multiple video clips and compiles them into a series of video files. In this step, the server integrates the generated video clips to create a time-coherent video file. It takes individual video clips as input, edits them, and generates a continuous video file as output.

[1057] Step 7:

[1058] The server uploads the generated video file to the cloud storage. In this step, the server saves the completed video file to the cloud storage via the Internet. It receives the completed video file as input and provides the cloud storage URL or access information as output.

[1059] Step 8:

[1060] Users access, download, and watch video files stored in cloud storage through a dedicated application. In this step, users access cloud storage using a dedicated user interface to view and save video files. Cloud storage access information is used as input, and video files are made available for viewing as output.

[1061] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1062] The present invention is a system that automatically generates individual videos for families with children and then uses an emotion engine to edit the videos based on the user's emotions. This system eliminates the need for parents to film and edit events, allowing them to easily record their children's memories and highlight moving moments. Specific embodiments of the system are described below.

[1063] System configuration

[1064] 1. High-resolution imaging equipment

[1065] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[1066] 2. Real-time video transmission method

[1067] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[1068] 3. Facial Recognition and Video Analysis

[1069] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1070] 4. How to track the subject

[1071] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[1072] 5. Automatic Clip Generation Method

[1073] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[1074] 6. Video file generation method

[1075] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[1076] 7. Cloud Storage Distribution Methods

[1077] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[1078] 8. User Interface

[1079] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[1080] Incorporating an emotion engine

[1081] 1. Emotion Engine

[1082] The server also has an emotion engine that recognizes the user's emotions from the video, analyzing emotions such as smile, sadness, and surprise in real time to obtain emotional data during the event.

[1083] 2. Emotional editing

[1084] The server then edits the video clips based on the recognized emotions, for example, prioritizing scenes with many smiling faces as highlight clips, or editing to emphasize moving scenes.

[1085] Specific examples

[1086] Emotion Recognition in the Case of Sports Day

[1087] 1. Event Settings

[1088] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[1089] 2. Start shooting

[1090] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[1091] 3. Video data transmission

[1092] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1093] 4. Emotion analysis

[1094] The server also uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[1095] 5. Clip Creation and Video Editing

[1096] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[1097] 6. Distribution of video files

[1098] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1099] This system allows users to record their children's growth and memories in a richer way, without the hassle of filming and editing during events, and enjoy footage that highlights moving moments. This increases the amount of time parents and children can spend together cheering and interacting with each other at the venue, allowing for more fulfilling parent-child time.

[1100] The processing flow will be explained below.

[1101] Step 1:

[1102] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[1103] Step 2:

[1104] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[1105] Step 3:

[1106] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[1107] Step 4:

[1108] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[1109] Step 5:

[1110] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1111] Step 6:

[1112] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[1113] Step 7:

[1114] The server also uses an emotion engine to analyze the emotions of the subjects in the video and the viewers in real time, detecting, for example, the smile or other emotions of a child crossing the finish line in a race.

[1115] Step 8:

[1116] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[1117] Step 9:

[1118] The server determines the importance of each clip based on the emotional data recognized by the emotion engine, and prioritizes highlights, such as scenes of smiling faces or cheering.

[1119] Step 10:

[1120] The server then edits the selected video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[1121] Step 11:

[1122] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[1123] Step 12:

[1124] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[1125] This series of processes eliminates the need for users to perform tedious filming and editing tasks during events, allowing them to spend more time enjoying their children's growth and memories. Furthermore, the emotion engine makes it possible to create videos that emphasize moving moments.

[1126] Example 2

[1127] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1128] For families with children, the time and effort required to film and edit events is a challenge. In particular, if parents are too focused on filming during an event, they lose the time they need to cheer on and interact with their children in real time. Furthermore, editing the footage to appropriately highlight moving moments and records of children's growth is difficult.

[1129] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1130] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject in the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for recognizing emotions in the video using an emotion engine, and means for editing the video based on the recognized emotions. This allows users to easily create videos that highlight their children's growth and moving moments without the hassle of filming and editing the event.

[1131] "High-resolution image capture device" refers to a capture device capable of capturing detailed footage of the entire event.

[1132] "Means for transmitting to a server in real time" refers to a communication means for immediately transferring captured video data to a server.

[1133] A "facial recognition algorithm" is software or a program that identifies the faces of people in a video and authenticates a specific individual.

[1134] "Subject" refers to an individual identified by a facial recognition algorithm.

[1135] A "tracking means" is a system for continuously recording the location of a subject recognized in the video.

[1136] The "means for automatically generating clips" is a function that automatically cuts out the portion of the video in which the subject appears.

[1137] The "means for generating a video file" is a system that combines multiple video clips and edits them into a single video file.

[1138] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.

[1139] "User interface" refers to the screen and input means used by users to operate software or a system.

[1140] An "emotion engine" is software or an algorithm that analyzes the emotions of people in a video and recognizes their emotional state.

[1141] "Means for editing video" refers to a function for processing captured video data and emphasizing or changing content based on specific scenes or emotions.

[1142] This system automatically generates individual videos for families with children and uses an emotion engine to edit the videos based on the user's emotions. This saves parents the trouble of filming and editing events, allowing them to easily record their children's memories and highlight moving moments.

[1143] This system is implemented using the following hardware and software.

[1144] High-resolution image capture device: The user installs a high-resolution camera at the event venue to capture the entire event. For example, the Sony Alpha series is used.

[1145] Real-time video transmission method: The device (e.g., laptop or Raspberry Pi) transmits the video data received from the camera to the server in real time. The communication method is a high-speed Internet connection (Wi-Fi or wired LAN).

[1146] Facial recognition algorithm: The server uses an AI facial recognition algorithm such as Amazon Rekognition to recognize the child's face from the received video data and identify the child by comparing it with previously uploaded facial image data.

[1147] Subject tracking method: The server continuously tracks the location information of the recognized subject within the video. It records the position in each frame, so it can follow the subject even if they move within the video.

[1148] Automatic clip generation: The server automatically extracts the video portion in which the recognized subject appears and generates individual video clips. The clips are then edited to be visually appealing, with automatic zooming and frame correction.

[1149] Video file generator: The server edits multiple video clips and compiles them into a series of video files, creating a consistent video of the entire event.

[1150] Cloud storage distribution method: The server uploads the completed video file to cloud storage (e.g., Google Drive). The data is saved in the cloud storage specified by the user.

[1151] User interface: Users can access cloud storage through a dedicated app (iOS or Android app) and download or watch video files.

[1152] Emotion engine: The server uses Amazon Rekognition and Microsoft Azure Emotion API to analyze the user's emotions in real time. It recognizes emotions such as smiles and surprise and records data during the event.

[1153] Emotion-based editing: The server edits video clips based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights and emphasizes moving scenes.

[1154] Example: Emotion recognition in the case of athletic meet

[1155] 1. Event Settings

[1156] Users set up a "Sports Day" event using a dedicated app, and upload facial images of the children involved in the event in advance.

[1157] 2. Start shooting

[1158] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[1159] 3. Video data transmission

[1160] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1161] 4. Emotion analysis

[1162] The server uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[1163] 5. Clip Creation and Video Editing

[1164] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[1165] 6. Distribution of video files

[1166] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated app.

[1167] Example prompts for generative AI models:

[1168] Perform emotion analysis on video data from a sports day and generate a video clip that emphasizes touching moments. The facial images of the target children and the actual video data can be found at the following links: (Facial image URL), (Video data URL). Edit the scenes in which smiling faces are detected as highlights.

[1169] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1170] Step 1:

[1171] The user sets up a high-resolution camera at the event venue, fixes the camera in a position where it can capture the entire event in detail, and starts recording, inputting video data of the entire event.

[1172] Step 2:

[1173] The device (laptop or Raspberry Pi) transmits the video data acquired from the camera to the server in real time. The communication method is a high-speed internet connection. The input in this step is the video data from the camera, and the output is the real-time video data transmission to the server.

[1174] Step 3:

[1175] The server analyzes the received video data using AI facial recognition algorithms such as Amazon Rekognition. It compares the data with previously uploaded facial image data and recognizes the child's face. The input for this step is real-time video data and facial image data, and the output is the location information of the recognized face.

[1176] Step 4:

[1177] The server tracks the recognized face in the video and keeps recording the subject's position information for each frame. The input is the position information of the recognized face, and the output is the subject's position data for each frame.

[1178] Step 5:

[1179] The server automatically extracts the portion of the video in which the recognized and tracked subject appears, generating individual video clips. The clips are then edited for visual clarity with automatic zoom and frame correction. The inputs are the subject's position data and video data, and the output is the video clip.

[1180] Step 6:

[1181] The server then edits the generated video clips into a series of video files. The clips are organized in chronological order to create a seamless, coherent video. The input is a series of video clips, and the output is an edited video file.

[1182] Step 7:

[1183] The server uploads the completed video file to the specified cloud storage, such as Google Drive or Amazon S3. The input is the edited video file, and the output is a notification that the upload to the cloud storage is complete.

[1184] Step 8:

[1185] Users use a dedicated app to access cloud storage and download or view video files. The interface is intuitive and easy to use. The input is the cloud storage URL, and the output is a video file that can be downloaded and viewed.

[1186] Step 9:

[1187] The server uses an emotion engine (such as Amazon Rekognition or Microsoft Azure Emotion API) to analyze the user's emotions in the video in real time. It recognizes emotions such as smiles and surprise and records emotional data during the event. The input is the video data, and the output is the emotion recognition results.

[1188] Step 10:

[1189] The server then edits the video clip based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights, and further emphasizes moving scenes. The input is the emotion recognition results and the video clip, and the output is an edited video based on the emotions.

[1190] (Application example 2)

[1191] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1192] In modern factory environments, worker safety management is a critical issue. Conventional monitoring methods have difficulty monitoring in real time whether workers are working in an appropriate state of health or under emotional stress. Furthermore, dangerous situations frequently occur at worksites due to the large number of machines in operation. Conventional safety monitoring systems rely on simple surveillance cameras, which are unable to analyze workers' emotions or fatigue levels. This makes it difficult to adequately ensure safety in the actual work environment. Therefore, there is a need for a system that can analyze workers' emotional state and fatigue levels in real time and respond quickly.

[1193] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video using a high-resolution image capture device, means for transmitting the video in real time, means for identifying a target person using a facial recognition algorithm, means for tracking the target person in the video, means for automatically generating clips, means for editing the clips to generate a video file, means for uploading the video file to cloud storage, means for providing a user interface for accessing the video file, means for analyzing emotions in the video in real time using an emotion engine and issuing a safety warning, and means for transmitting the analysis data in real time, and means for storing the data in cloud storage. This makes it possible to monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[1194] A "high resolution imaging device" is an imaging device for capturing highly detailed videos and images.

[1195] "Means for transmitting to the server in real time" refers to the communication technology for instantly sending captured video data to the server.

[1196] A "face recognition algorithm" is a computer program that identifies faces in video and compares them with pre-registered facial information.

[1197] "Means for tracking a subject within a video" refers to a technology that tracks the position of a specific person even if they move within the video.

[1198] "Means for automatically generating clips" refers to technology that automatically cuts out specific video segments and generates them as short video clips.

[1199] The "means for generating a video file" refers to the technology for editing clips and putting them together as a series of videos.

[1200] "Means for uploading to cloud storage" refers to a technology for storing the generated video file in a remote storage area on the Internet.

[1201] A "user interface" is the portion of software that provides the visual and operational elements through which a user interacts with a system.

[1202] An "emotion engine" is an algorithm and software that analyzes the emotions of people in a video and identifies their emotional state.

[1203] "Means for issuing safety warnings" refers to technology for issuing warnings when dangerous situations are detected as a result of emotion analysis.

[1204] "Data transmission means" refers to a communication technique for sending the analyzed data to another location.

[1205] "Means by which data is stored in cloud storage" refers to the technology used to store analyzed data and generated videos in the cloud.

[1206] The present invention relates to a worker safety monitoring system for a factory, which uses a high-resolution image capture device, a high-performance server, and cloud storage. The system analyzes the facial expressions and postures of workers working in the factory and can perform safety management in real time. Specific embodiments of the system are described below.

[1207] System configuration

[1208] 1. High-resolution imaging equipment

[1209] The user installs a high-resolution camera in the factory, which covers a wide area and captures detailed images of workers. For example, the Logitech C920 is used as this camera.

[1210] 2. Real-time video transmission method

[1211] The device transmits the captured video data in real time to a server using an internet connection, where the video data is processed immediately.

[1212] 3. Facial Recognition Algorithm

[1213] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the worker's face. At this time, it compares it with pre-registered facial images to identify the specific worker. OpenCV is used for facial recognition.

[1214] 4. How to track the subject

[1215] The server tracks the recognized worker's face in the video and records the position information in each frame, making it possible to track the target's position even if the target moves within the video.

[1216] 5. Emotion analysis

[1217] The server uses an emotion engine to analyze the emotions of workers in the video in real time. For example, it analyzes emotions such as smiles, sadness, and surprise, and generates data useful for safety management. An AI model using TensorFlow / Keras is used for emotion analysis.

[1218] 6. Issuance of safety warnings

[1219] Based on the analysis results of the emotion engine, the server quickly issues safety warnings, including visual alerts and audio notifications, if a worker is in a dangerous or overly fatigued state.

[1220] 7. Data transmission and storage in cloud storage

[1221] The server uploads the analyzed data and generated video clips to cloud storage and also transmits the analyzed data to other devices in real time, allowing factory managers to monitor the safety status of workers even from remote locations.

[1222] 8. User Interface

[1223] Users can access cloud storage and check saved data and video clips using a dedicated application. The intuitive user interface allows users to quickly obtain the information they need.

[1224] Specific examples

[1225] To implement a system to monitor the safety of workers in a factory, high-resolution cameras are installed in each work area. The cameras capture the facial expressions and postures of workers in real time and send the footage to a server. The server then uses a facial recognition algorithm and an emotion engine to identify the worker and perform emotion analysis. Based on the analysis results, a safety warning is issued immediately if the worker is in danger. Furthermore, the real-time analysis data is stored in cloud storage and can be accessed by managers through a dedicated application.

[1226] Example prompts for generative AI models

[1227] Design a system that uses facial recognition and emotion analysis to monitor the safety of workers in a factory. A high-resolution camera captures images of workers working and analyzes them in real time using an emotion engine. If a worker is in danger or overly fatigued, an alert will be issued and the data will be stored in the cloud. Please provide a concrete code example.

[1228] As a result, a system can be provided that can monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[1229] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1230] Step 1:

[1231] Users install high-resolution cameras in the work area of ​​their factories, which capture detailed images of workers' faces and the work they are doing.

[1232] Input: spatial information (camera installation position), high-resolution camera

[1233] Output: Real-time video data

[1234] Step 2:

[1235] The device transmits the captured real-time video data to a server via the Internet, where streaming technology is used to efficiently transmit large amounts of data.

[1236] Input: Real-time video data

[1237] Output: Video data sent to the server

[1238] Step 3:

[1239] The server uses OpenCV to analyze the received video data and identifies the worker's face using a facial recognition algorithm, which then compares it with pre-registered facial information to identify each worker.

[1240] Input: Transmitted video data, pre-registered face data

[1241] Output: Video data including identified workers (with location information)

[1242] Step 4:

[1243] The server tracks the face of the identified worker and records their position in each frame, allowing the worker to be tracked in real time even if they move within the camera's field of view.

[1244] Input: Video data containing identified workers

[1245] Output: Video data with location information of tracked workers

[1246] Step 5:

[1247] The server automatically generates a clip of the worker's face, zooming and adjusting the frame as needed, using an AI model to generate the optimal clip.

[1248] Input: Location-based video data of tracked workers

[1249] Output: Automatically generated clip video data

[1250] Step 6:

[1251] The server uses an emotion engine to analyze the emotions of the workers in the video clips, specifically identifying emotional states such as smiling or surprised using TensorFlow / Keras.

[1252] Input: Clip video data

[1253] Output: Parsed emotion data (with emotion labels)

[1254] Step 7:

[1255] The server then issues safety alerts based on the analyzed emotion data, for example, visual and audio alerts if a worker is in a dangerous emotional state (anger or sadness).

[1256] Input: Emotion-labeled analysis data

[1257] Output: Safety warning (alert)

[1258] Step 8:

[1259] The server uploads the analyzed data and generated video clips to cloud storage, and also transmits the analyzed data in real time to a remote administrator terminal.

[1260] Input: Analysis data, clip video data

[1261] Output: Data stored in the cloud, analysis data sent to the administrator's terminal

[1262] Step 9:

[1263] Users (administrators) can access cloud storage and check stored data and video clips using a dedicated application. The application has an intuitive interface, allowing users to quickly obtain the information they need.

[1264] Input: Data stored in the cloud (video clips, analysis data)

[1265] Output: Monitoring information available to administrators (interface operations)

[1266] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1267] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1268] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1269] [Fourth embodiment]

[1270] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1271] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1272] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1273] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1274] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1275] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1276] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1277] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1278] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1279] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1280] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1281] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1282] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1283] The present invention is a system for automatically generating personalized videos for families with children, allowing parents to easily record their children's memories without the need to film and edit events. This system is implemented as follows.

[1284] System configuration

[1285] 1. High-resolution imaging equipment

[1286] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[1287] 2. Real-time video transmission method

[1288] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[1289] 3. Facial Recognition and Video Analysis

[1290] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1291] 4. How to track the subject

[1292] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[1293] 5. Automatic Clip Generation Method

[1294] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[1295] 6. Video file generation method

[1296] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[1297] 7. Cloud Storage Distribution Methods

[1298] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[1299] 8. User Interface

[1300] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[1301] Specific examples

[1302] Sports day case

[1303] 1. Event Settings

[1304] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[1305] 2. Start shooting

[1306] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[1307] 3. Video data transmission

[1308] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1309] 4. Clip Creation and Video Editing

[1310] The server automatically extracts and edits the footage of the child in question into a clip. For example, a clip of the child running is clipped and edited to include all the necessary scenes.

[1311] 5. Distribution of video files

[1312] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1313] This system allows parents to easily record and enjoy their children's growth and memories without the hassle of filming and editing during events, increasing the amount of time parents can spend cheering and interacting with their children at the venue and creating a more fulfilling parent-child experience.

[1314] The processing flow will be explained below.

[1315] Step 1:

[1316] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[1317] Step 2:

[1318] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[1319] Step 3:

[1320] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[1321] Step 4:

[1322] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[1323] Step 5:

[1324] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1325] Step 6:

[1326] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[1327] Step 7:

[1328] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[1329] Step 8:

[1330] The server then edits the multiple video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[1331] Step 9:

[1332] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[1333] Step 10:

[1334] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[1335] This series of processes eliminates the need for the user to perform tedious shooting and editing tasks during the event, allowing them to have more time to enjoy their children's growth and memories.

[1336] Example 1

[1337] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1338] Nowadays, parents spend a lot of time and effort recording their children's growth and important events. Filming and editing events is particularly tedious work, requiring parents to spend a lot of time on the task, which can result in parents being unable to fully enjoy the event itself. Furthermore, the quality of the footage recorded varies depending on the cameraman's skill and the equipment used. Therefore, there is a need for a system that allows parents to easily obtain high-quality footage.

[1339] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1340] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video in real time to a data processing device, means for identifying a subject using a facial recognition algorithm on the data processing device, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for automatically zooming and frame correction for the subject, and means for pre-registering facial images of the subject. This allows parents to easily record and enjoy high-quality video of their children's growth and activities during the event without the effort of filming or editing.

[1341] "High-resolution image capturing device" refers to any imaging device that has a high pixel count and can capture detailed, clear images.

[1342] "Means for transmitting video in real time" refers to communication technology and devices for transferring captured video data to a server or data processing device immediately without delay.

[1343] A "facial recognition algorithm" refers to software or computational methods used to identify faces in images and detect specific individuals.

[1344] "Means for identifying subjects" refers to functions and technologies for recognizing and identifying individual people within video footage.

[1345] "Means for tracking a subject within a video stream" refers to the technology or computational methods used to continuously track the location of an identified person as they move within the video stream.

[1346] "Means for automatically generating clips" refers to functions and technologies for cutting out footage showing specific people and automatically editing it into short video segments.

[1347] "Means for generating individual video files" refers to the functions and technologies for editing and integrating multiple video clips into a continuous video file.

[1348] "Means for uploading to cloud storage" refers to the functionality and technology for storing the generated video files in a cloud-based storage device via the Internet.

[1349] "Means for providing a user interface" refers to display screens and operating means that make it easier for users to operate systems and services.

[1350] "Means for automatic zoom and frame correction" refers to functions and technologies for zooming in on a specific subject in a video or automatically adjusting the composition of the video.

[1351] "Means for pre-registering facial images" refers to functions and technologies that allow the system to load and store facial image data of the subject in advance.

[1352] MODE FOR CARRYING OUT THE INVENTION

[1353] This invention is an automatic individual video generation system designed for families with children, which reduces the effort required for parents to film and edit events and allows them to easily record their children's growth and memories. This system is implemented using the following configuration and process.

[1354] System configuration

[1355] 1. High-resolution imaging equipment

[1356] Users install a high-resolution camera at the event venue. This camera takes still pictures and captures the entire event in detail. This device can be, for example, a 4K video camera or a high-performance digital SLR camera.

[1357] 2. Real-time video transmission

[1358] The device transmits the captured video data to a server in real time. This process uses an internet connection, so the data is sent to the server instantly. Specifically, a smartphone or tablet receives the data from the camera and sends it to the server via Wi-Fi or mobile network.

[1359] 3. Facial Recognition and Video Analysis

[1360] The server analyzes the received video data and uses AI facial recognition algorithms to identify the specific child's face, using facial recognition software such as Amazon Rekognition or OpenCV.

[1361] 4. Subject Tracking

[1362] The server tracks the recognized child's face in the video and records its position in each frame, allowing it to track the child's position even if the child moves within the video.

[1363] 5. Automatic generation of video clips

[1364] The server automatically extracts the footage containing the subject and generates individual video clips using video editing software such as FFmpeg, and automatically applies zoom and frame correction to create visually pleasing clips.

[1365] 6. Generating video files

[1366] The server then assembles the edited video clips into a series of video files, a process that organizes the clips chronologically to create a coherent video.

[1367] 7. Upload to cloud storage

[1368] The server then uploads the generated video files to cloud storage, typically a cloud-based storage service such as Google Drive or Dropbox.

[1369] 8. Providing a User Interface

[1370] Users access video files stored in cloud storage through a dedicated application, which is intuitive and easy to use, designed to allow users to easily find, download, and watch videos.

[1371] Specific examples

[1372] In the case of sports day

[1373] 1. Event setup: The user sets up the "Sports Day" event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[1374] 2. Start shooting: The user sets up a high-resolution camera at the sports day venue and starts shooting from a position that covers the entire venue.

[1375] 3. Real-time video transmission: The device transmits video data to the server in real time. The server analyzes the received video and performs facial recognition and tracking of the subject.

[1376] 4. Clip generation and video editing: The server automatically extracts and edits the video clips containing the target children. For example, it clips the footage of the children running and edits it to include all the necessary scenes.

[1377] 5. Video file distribution: The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1378] Prompt Sentence Examples

[1379] markdown

[1380] Describe a system that automatically edits and saves footage of children's sports events to cloud storage. The user installs a high-resolution camera and configures the event using a dedicated app. The device transmits the footage in real time to a server, which uses AI facial recognition to identify the children and automatically edits the footage. The final video file is uploaded to cloud storage and can be accessed by the user.

[1381] This system allows parents to record and enjoy high-quality footage of their children's growth and memories without the effort of filming and editing during events. By increasing the amount of time parents spend cheering and interacting with their children at the venue, parents can spend more fulfilling time with their children.

[1382] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1383] Step 1:

[1384] The user launches a dedicated application and sets up the event. First, they move to the settings screen within the application and enter the event name (e.g., "Sports Day") and date. Next, they upload a facial image of the child in question. This facial image will be used for facial recognition later. The event information and facial image set by the user are entered. The entered data is saved within the application.

[1385] Step 2:

[1386] Users set up high-resolution cameras at the event venue. They position the cameras so that they can cover the entire event, and adjust the angle and height as needed. When they're ready to start recording, they press the camera's record button to begin recording. The cameras generate high-quality video data in real time, allowing the entire event to be recorded in detail.

[1387] Step 3:

[1388] The terminal (device connected to the camera) compresses the captured video data in real time and sends it to a server via the Internet. Video data is input and output in a compressed form. The video data sent from the terminal is received by the server. This process achieves efficient data transfer.

[1389] Step 4:

[1390] The server analyzes the received video data and uses an AI facial recognition algorithm (such as Amazon Rekognition or OpenCV) to match it with previously uploaded facial images. The facial image is input and the face of a specific child is recognized. The server identifies the target person based on this processing. The location of the specific child within the video is confirmed through facial recognition.

[1391] Step 5:

[1392] The server tracks the recognized subject's face in the video and records their location in each frame. The location information obtained through facial recognition is input, and tracking data including the subject's location is output. The server can use this tracking data to track the location of a specific child over time.

[1393] Step 6:

[1394] The server automatically extracts the portion of the video that shows the subject and generates a separate video clip. This process is performed using video editing software (e.g., FFmpeg). Tracking data is input and a video clip with the subject at the center is output. The server then applies zoom and frame correction to this clip to make it more visually appealing.

[1395] Step 7:

[1396] The server chronologically organizes multiple video clips into a series of video files. Edited clips are input and a coherent video file is output. The server encodes this video file and converts it to the optimal format (e.g., MP4).

[1397] Step 8:

[1398] The server uploads the generated video file to the cloud storage specified by the user (e.g., Google Drive or Dropbox). The completed video file is input and saved in the cloud storage. The upload process sends the file over the Internet and stores it securely in the cloud.

[1399] Step 9:

[1400] Users access cloud storage through a dedicated application to view uploaded video files. Users can search for video files and download or stream them. The application provides an intuitive interface, making it easy to find the video you want.

[1401] (Application example 1)

[1402] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1403] In conventional virtual shopping, parents must manually record and then edit videos to record their children's real-time reactions. This process requires time and effort for parents, hindering an intuitive and natural shopping experience. Furthermore, manual filming and editing increases the risk of missing important moments. Therefore, there is a need for a system that can automatically capture and record children's reactions and fun moments.

[1404] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1405] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject within the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, and means for automatically capturing and recording a child's reactions and moments of enjoyment during virtual shopping. This eliminates the need for parents to manually shoot and edit videos, and allows parents to easily record important moments during virtual shopping without missing them.

[1406] A "high resolution image capture device" is a camera device capable of capturing video at high resolution.

[1407] "Means for transmitting to the server in real time" refers to technology that transfers video data to the server immediately without delay.

[1408] "Means for identifying a subject using a facial recognition algorithm on a server" refers to a method in which a server uses facial recognition technology to identify a specific person.

[1409] "Means for tracking a subject in a video" refers to a technology that continuously tracks the position and movements of a specific person in a video.

[1410] "Means for automatically generating clips" refers to a technology that automatically cuts out specific parts of video and edits them into short video segments.

[1411] The "means for generating a video file" refers to a method for integrating multiple video clips and editing them into a single continuous video file.

[1412] "Means for uploading to cloud storage" refers to a technology that stores the generated video files on a remote server and makes them accessible via the Internet.

[1413] "Means for providing a user interface" refers to operation screens and applications that allow users to interact with the system intuitively.

[1414] "Virtual shopping" is the act of selecting and purchasing products in a virtual space using the Internet.

[1415] "A means for automatically capturing and recording children's reactions and moments of enjoyment" is a technology that detects children's facial expressions and behavior in real time and automatically records those moments as video.

[1416] This invention is a system that automatically captures and records moments that parents enjoy with their children during virtual shopping, allowing parents to preserve memories without any hassle.

[1417] System configuration

[1418] 1. High-resolution imaging equipment

[1419] Users place a high-resolution camera in the virtual shopping environment to capture detailed, high-resolution footage of the entire shopping experience. The camera remains stationary and captures footage of the shopping experience.

[1420] 2. Real-time video transmission method

[1421] The device (smartphone or head-mounted display) transmits the captured video data to a server in real time using an internet connection, and video data captured while shopping is sent to the server instantly.

[1422] 3. Facial Recognition and Video Analysis

[1423] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1424] 4. How to track the subject

[1425] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[1426] 5. Automatic Clip Generation Method

[1427] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[1428] 6. Video file generation method

[1429] The server compiles the video clips into a series of video files, which are organized chronologically to create a coherent video of the entire shopping experience.

[1430] 7. Cloud Storage Distribution Methods

[1431] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[1432] 8. User Interface

[1433] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[1434] Hardware and software used

[1435] Hardware

[1436] Webcam

[1437] Smartphone

[1438] head-mounted display

[1439] software

[1440] OpenCV (image processing library)

[1441] dlib (face recognition algorithm)

[1442] moviepy (video editing library)

[1443] Specific application examples

[1444] The case for virtual shopping

[1445] 1. Pre-registration of facial images

[1446] Users upload a photo of their child's face in advance through a dedicated application.

[1447] 2. Start shooting

[1448] Users place high-resolution cameras in the virtual shopping environment and capture footage in real time.

[1449] 3. Video data transmission

[1450] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1451] 4. Clip Creation and Video Editing

[1452] The server automatically cuts out clips of footage in which the child appears, and edits them together, for example, when the child shows interest in a product, and edits the clips to include all the important moments.

[1453] 5. Distribution of video files

[1454] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1455] Prompt Sentence Examples

[1456] "Develop an application that automatically captures a child's smile or surprised expression while they are shopping, and creates a video of their memories that can be viewed later."

[1457] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1458] Step 1:

[1459] Users first upload a facial image of their child through a dedicated application. In this step, users register their child's facial image data in the system using a smartphone or computer. The system receives this image data as input and generates and stores facial feature data for use in the facial recognition algorithm.

[1460] Step 2:

[1461] The user installs a high-resolution camera in the virtual shopping environment and captures video in real time. The video captured by the camera is positioned to cover the entire view of the shop. The camera captures video data in real time and transmits it to the terminal. The terminal then transmits the received video data to the server in real time.

[1462] Step 3:

[1463] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face. In this step, the server receives real-time video data as input and runs the facial recognition algorithm. The algorithm compares the face in the real-time video with pre-registered facial feature data, and if there is a match, the child's face is identified.

[1464] Step 4:

[1465] The server tracks the recognized subject's face in the video and records its position in each frame. In this step, the server continues to record the position of the identified face in each frame of the video. For each frame received as input, it provides the face position information as output.

[1466] Step 5:

[1467] The server automatically extracts the video portion where the subject appears and generates a separate video clip. In this step, the server analyzes the video based on the facial position information, extracts the portion where the subject appears, and creates a clip. Zoom and frame correction are also applied automatically. The input is the tracked video frame and position information, and the output is a visually pleasing video clip.

[1468] Step 6:

[1469] The server edits multiple video clips and compiles them into a series of video files. In this step, the server integrates the generated video clips to create a time-coherent video file. It takes individual video clips as input, edits them, and generates a continuous video file as output.

[1470] Step 7:

[1471] The server uploads the generated video file to the cloud storage. In this step, the server saves the completed video file to the cloud storage via the Internet. It receives the completed video file as input and provides the cloud storage URL or access information as output.

[1472] Step 8:

[1473] Users access, download, and watch video files stored in cloud storage through a dedicated application. In this step, users access cloud storage using a dedicated user interface to view and save video files. Cloud storage access information is used as input, and video files are made available for viewing as output.

[1474] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1475] The present invention is a system that automatically generates individual videos for families with children and then uses an emotion engine to edit the videos based on the user's emotions. This system eliminates the need for parents to film and edit events, allowing them to easily record their children's memories and highlight moving moments. Specific embodiments of the system are described below.

[1476] System configuration

[1477] 1. High-resolution imaging equipment

[1478] The user installs a high-resolution camera at the event venue to capture detailed footage of the entire event. The camera remains stationary and captures footage of the entire event.

[1479] 2. Real-time video transmission method

[1480] The device transmits the captured video data in real time to a server using an internet connection, allowing the event video data to be processed immediately.

[1481] 3. Facial Recognition and Video Analysis

[1482] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1483] 4. How to track the subject

[1484] The server tracks the recognized subject's face in the video and records their position in each frame, making it possible to track their position even if they move within the video.

[1485] 5. Automatic Clip Generation Method

[1486] The server automatically extracts the footage containing the subject and generates individual video clips, which are then automatically zoomed, frame-corrected, and edited for optimal viewing.

[1487] 6. Video file generation method

[1488] The server compiles the multiple video clips into a series of video files, which are organized chronologically to create a coherent video of the entire event.

[1489] 7. Cloud Storage Distribution Methods

[1490] The server uploads the generated video file to a cloud storage designated by the user and accessible via the Internet.

[1491] 8. User Interface

[1492] Users can access, download, and watch video files stored in cloud storage through a dedicated app. The user interface is intuitive and easy to use, making it easy to find the video you want.

[1493] Incorporating an emotion engine

[1494] 1. Emotion Engine

[1495] The server also has an emotion engine that recognizes the user's emotions from the video, analyzing emotions such as smile, sadness, and surprise in real time to obtain emotional data during the event.

[1496] 2. Emotional editing

[1497] The server then edits the video clips based on the recognized emotions, for example, prioritizing scenes with many smiling faces as highlight clips, or editing to emphasize moving scenes.

[1498] Specific examples

[1499] Emotion Recognition in the Case of Sports Day

[1500] 1. Event Settings

[1501] The user sets up a sports day event using a dedicated application, and uploads facial images of the children involved in the event in advance.

[1502] 2. Start shooting

[1503] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[1504] 3. Video data transmission

[1505] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1506] 4. Emotion analysis

[1507] The server also uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[1508] 5. Clip Creation and Video Editing

[1509] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[1510] 6. Distribution of video files

[1511] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated application.

[1512] This system allows users to record their children's growth and memories in a richer way, without the hassle of filming and editing during events, and enjoy footage that highlights moving moments. This increases the amount of time parents and children can spend together cheering and interacting with each other at the venue, allowing for more fulfilling parent-child time.

[1513] The processing flow will be explained below.

[1514] Step 1:

[1515] The user enters detailed event information (date, time, location, and facial image of the target child) through a dedicated application. For example, the user can set information such as "October 10, 2023, XX Elementary School Sports Day, child's name."

[1516] Step 2:

[1517] The server receives the input event information and registers it in a database, along with a facial image of the child in question.

[1518] Step 3:

[1519] The user sets up a high-resolution camera at the event venue and starts recording. The camera is fixed so that it captures the entire event.

[1520] Step 4:

[1521] The device transmits the captured video data to a server in real time via an internet connection, with the data being sent in streaming format.

[1522] Step 5:

[1523] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the child's face, which is then compared with pre-registered facial images to identify the specific child.

[1524] Step 6:

[1525] The server tracks the subject within the video based on the coordinate information for each frame in which a face is recognized. All video frames during this time are analyzed to sequentially track the subject's movements.

[1526] Step 7:

[1527] The server also uses an emotion engine to analyze the emotions of the subjects in the video and the viewers in real time, detecting, for example, the smile or other emotions of a child crossing the finish line in a race.

[1528] Step 8:

[1529] The server automatically extracts the video footage showing the subject and generates individual clips, which are then edited for visual clarity using automatic zooming and frame correction.

[1530] Step 9:

[1531] The server determines the importance of each clip based on the emotional data recognized by the emotion engine, and prioritizes highlights, such as scenes of smiling faces or cheering.

[1532] Step 10:

[1533] The server then edits the selected video clips and compiles them into a series of video files, adjusting the transitions between the clips and the audio to create a continuous video.

[1534] Step 11:

[1535] The server then uploads the completed video file to the cloud storage service specified by the user. This process is done automatically, with no special user action required.

[1536] Step 12:

[1537] Users can access, download, and watch video files stored in cloud storage through a dedicated application. Uploaded videos are displayed as thumbnails using the preview function, making viewing easy.

[1538] This series of processes eliminates the need for users to perform tedious filming and editing tasks during events, allowing them to spend more time enjoying their children's growth and memories. Furthermore, the emotion engine makes it possible to create videos that emphasize moving moments.

[1539] Example 2

[1540] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1541] For families with children, the time and effort required to film and edit events is a challenge. In particular, if parents are too focused on filming during an event, they lose the time they need to cheer on and interact with their children in real time. Furthermore, editing the footage to appropriately highlight moving moments and records of children's growth is difficult.

[1542] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1543] In this invention, the server includes means for capturing video of the entire event using a high-resolution image capture device, means for transmitting the video to the server in real time, means for identifying a subject using a facial recognition algorithm on the server, means for tracking the subject in the video, means for automatically generating clips focusing on the subject, means for editing the clips to generate individual video files, means for uploading the generated video files to cloud storage, means for providing a user interface for accessing the video files, means for recognizing emotions in the video using an emotion engine, and means for editing the video based on the recognized emotions. This allows users to easily create videos that highlight their children's growth and moving moments without the hassle of filming and editing the event.

[1544] "High-resolution image capture device" refers to a capture device capable of capturing detailed footage of the entire event.

[1545] "Means for transmitting to a server in real time" refers to a communication means for immediately transferring captured video data to a server.

[1546] A "facial recognition algorithm" is software or a program that identifies the faces of people in a video and authenticates a specific individual.

[1547] "Subject" refers to an individual identified by a facial recognition algorithm.

[1548] A "tracking means" is a system for continuously recording the location of a subject recognized in the video.

[1549] The "means for automatically generating clips" is a function that automatically cuts out the portion of the video in which the subject appears.

[1550] The "means for generating a video file" is a system that combines multiple video clips and edits them into a single video file.

[1551] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.

[1552] "User interface" refers to the screen and input means used by users to operate software or a system.

[1553] An "emotion engine" is software or an algorithm that analyzes the emotions of people in a video and recognizes their emotional state.

[1554] "Means for editing video" refers to a function for processing captured video data and emphasizing or changing content based on specific scenes or emotions.

[1555] This system automatically generates individual videos for families with children and uses an emotion engine to edit the videos based on the user's emotions. This saves parents the trouble of filming and editing events, allowing them to easily record their children's memories and highlight moving moments.

[1556] This system is implemented using the following hardware and software.

[1557] High-resolution image capture device: The user installs a high-resolution camera at the event venue to capture the entire event. For example, the Sony Alpha series is used.

[1558] Real-time video transmission method: The device (e.g., laptop or Raspberry Pi) transmits the video data received from the camera to the server in real time. The communication method is a high-speed Internet connection (Wi-Fi or wired LAN).

[1559] Facial recognition algorithm: The server uses an AI facial recognition algorithm such as Amazon Rekognition to recognize the child's face from the received video data and identify the child by comparing it with previously uploaded facial image data.

[1560] Subject tracking method: The server continuously tracks the location information of the recognized subject within the video. It records the position in each frame, so it can follow the subject even if they move within the video.

[1561] Automatic clip generation: The server automatically extracts the video portion in which the recognized subject appears and generates individual video clips. The clips are then edited to be visually appealing, with automatic zooming and frame correction.

[1562] Video file generator: The server edits multiple video clips and compiles them into a series of video files, creating a consistent video of the entire event.

[1563] Cloud storage distribution method: The server uploads the completed video file to cloud storage (e.g., Google Drive). The data is saved in the cloud storage specified by the user.

[1564] User interface: Users can access cloud storage through a dedicated app (iOS or Android app) and download or watch video files.

[1565] Emotion engine: The server uses Amazon Rekognition and Microsoft Azure Emotion API to analyze the user's emotions in real time. It recognizes emotions such as smiles and surprise and records data during the event.

[1566] Emotion-based editing: The server edits video clips based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights and emphasizes moving scenes.

[1567] Example: Emotion recognition in the case of athletic meet

[1568] 1. Event Settings

[1569] Users set up a "Sports Day" event using a dedicated app, and upload facial images of the children involved in the event in advance.

[1570] 2. Start shooting

[1571] The user sets up a high-resolution camera at the sports day venue and starts recording from a position that covers the entire venue.

[1572] 3. Video data transmission

[1573] The device sends the captured video in real time to a server, which then analyzes the video data and performs facial recognition and subject tracking.

[1574] 4. Emotion analysis

[1575] The server uses an emotion engine to recognize the emotions of viewers and children. For example, if a smile is detected at the moment a child crosses the finish line, the server will record that scene in a special way to highlight it.

[1576] 5. Clip Creation and Video Editing

[1577] The server automatically cuts out the video clips showing the target child and edits them. It also selects highlight scenes from the video based on the results of emotion analysis, creating a moving edit.

[1578] 6. Distribution of video files

[1579] The server uploads the edited video file to cloud storage, allowing users to access the video through a dedicated app.

[1580] Example prompts for generative AI models:

[1581] Perform emotion analysis on video data from a sports day and generate a video clip that emphasizes touching moments. The facial images of the target children and the actual video data can be found at the following links: (Facial image URL), (Video data URL). Edit the scenes in which smiling faces are detected as highlights.

[1582] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1583] Step 1:

[1584] The user sets up a high-resolution camera at the event venue, fixes the camera in a position where it can capture the entire event in detail, and starts recording, inputting video data of the entire event.

[1585] Step 2:

[1586] The device (laptop or Raspberry Pi) transmits the video data acquired from the camera to the server in real time. The communication method is a high-speed internet connection. The input in this step is the video data from the camera, and the output is the real-time video data transmission to the server.

[1587] Step 3:

[1588] The server analyzes the received video data using AI facial recognition algorithms such as Amazon Rekognition. It compares the data with previously uploaded facial image data and recognizes the child's face. The input for this step is real-time video data and facial image data, and the output is the location information of the recognized face.

[1589] Step 4:

[1590] The server tracks the recognized face in the video and keeps recording the subject's position information for each frame. The input is the position information of the recognized face, and the output is the subject's position data for each frame.

[1591] Step 5:

[1592] The server automatically extracts the portion of the video in which the recognized and tracked subject appears, generating individual video clips. The clips are then edited for visual clarity with automatic zoom and frame correction. The inputs are the subject's position data and video data, and the output is the video clip.

[1593] Step 6:

[1594] The server then edits the generated video clips into a series of video files. The clips are organized in chronological order to create a seamless, coherent video. The input is a series of video clips, and the output is an edited video file.

[1595] Step 7:

[1596] The server uploads the completed video file to the specified cloud storage, such as Google Drive or Amazon S3. The input is the edited video file, and the output is a notification that the upload to the cloud storage is complete.

[1597] Step 8:

[1598] Users use a dedicated app to access cloud storage and download or view video files. The interface is intuitive and easy to use. The input is the cloud storage URL, and the output is a video file that can be downloaded and viewed.

[1599] Step 9:

[1600] The server uses an emotion engine (such as Amazon Rekognition or Microsoft Azure Emotion API) to analyze the user's emotions in the video in real time. It recognizes emotions such as smiles and surprise and records emotional data during the event. The input is the video data, and the output is the emotion recognition results.

[1601] Step 10:

[1602] The server then edits the video clip based on the recognized emotions. For example, it prioritizes scenes with many smiling faces as highlights, and further emphasizes moving scenes. The input is the emotion recognition results and the video clip, and the output is an edited video based on the emotions.

[1603] (Application example 2)

[1604] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1605] In modern factory environments, worker safety management is a critical issue. Conventional monitoring methods have difficulty monitoring in real time whether workers are working in an appropriate state of health or under emotional stress. Furthermore, dangerous situations frequently occur at worksites due to the large number of machines in operation. Conventional safety monitoring systems rely on simple surveillance cameras, which are unable to analyze workers' emotions or fatigue levels. This makes it difficult to adequately ensure safety in the actual work environment. Therefore, there is a need for a system that can analyze workers' emotional state and fatigue levels in real time and respond quickly.

[1606] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video using a high-resolution image capture device, means for transmitting the video in real time, means for identifying a target person using a facial recognition algorithm, means for tracking the target person in the video, means for automatically generating clips, means for editing the clips to generate a video file, means for uploading the video file to cloud storage, means for providing a user interface for accessing the video file, means for analyzing emotions in the video in real time using an emotion engine and issuing a safety warning, and means for transmitting the analysis data in real time, and means for storing the data in cloud storage. This makes it possible to monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[1607] A "high resolution imaging device" is an imaging device for capturing highly detailed videos and images.

[1608] "Means for transmitting to the server in real time" refers to the communication technology for instantly sending captured video data to the server.

[1609] A "face recognition algorithm" is a computer program that identifies faces in video and compares them with pre-registered facial information.

[1610] "Means for tracking a subject within a video" refers to a technology that tracks the position of a specific person even if they move within the video.

[1611] "Means for automatically generating clips" refers to technology that automatically cuts out specific video segments and generates them as short video clips.

[1612] The "means for generating a video file" refers to the technology for editing clips and putting them together as a series of videos.

[1613] "Means for uploading to cloud storage" refers to a technology for storing the generated video file in a remote storage area on the Internet.

[1614] A "user interface" is the portion of software that provides the visual and operational elements through which a user interacts with a system.

[1615] An "emotion engine" is an algorithm and software that analyzes the emotions of people in a video and identifies their emotional state.

[1616] "Means for issuing safety warnings" refers to technology for issuing warnings when dangerous situations are detected as a result of emotion analysis.

[1617] "Data transmission means" refers to a communication technique for sending the analyzed data to another location.

[1618] "Means by which data is stored in cloud storage" refers to the technology used to store analyzed data and generated videos in the cloud.

[1619] The present invention relates to a worker safety monitoring system for a factory, which uses a high-resolution image capture device, a high-performance server, and cloud storage. The system analyzes the facial expressions and postures of workers working in the factory and can perform safety management in real time. Specific embodiments of the system are described below.

[1620] System configuration

[1621] 1. High-resolution imaging equipment

[1622] The user installs a high-resolution camera in the factory, which covers a wide area and captures detailed images of workers. For example, the Logitech C920 is used as this camera.

[1623] 2. Real-time video transmission method

[1624] The device transmits the captured video data in real time to a server using an internet connection, where the video data is processed immediately.

[1625] 3. Facial Recognition Algorithm

[1626] The server analyzes the received video data and uses an AI facial recognition algorithm to recognize the worker's face. At this time, it compares it with pre-registered facial images to identify the specific worker. OpenCV is used for facial recognition.

[1627] 4. How to track the subject

[1628] The server tracks the recognized worker's face in the video and records the position information in each frame, making it possible to track the target's position even if the target moves within the video.

[1629] 5. Emotion analysis

[1630] The server uses an emotion engine to analyze the emotions of workers in the video in real time. For example, it analyzes emotions such as smiles, sadness, and surprise, and generates data useful for safety management. An AI model using TensorFlow / Keras is used for emotion analysis.

[1631] 6. Issuance of safety warnings

[1632] Based on the analysis results of the emotion engine, the server quickly issues safety warnings, including visual alerts and audio notifications, if a worker is in a dangerous or overly fatigued state.

[1633] 7. Data transmission and storage in cloud storage

[1634] The server uploads the analyzed data and generated video clips to cloud storage and also transmits the analyzed data to other devices in real time, allowing factory managers to monitor the safety status of workers even from remote locations.

[1635] 8. User Interface

[1636] Users can access cloud storage and check saved data and video clips using a dedicated application. The intuitive user interface allows users to quickly obtain the information they need.

[1637] Specific examples

[1638] To implement a system to monitor the safety of workers in a factory, high-resolution cameras are installed in each work area. The cameras capture the facial expressions and postures of workers in real time and send the footage to a server. The server then uses a facial recognition algorithm and an emotion engine to identify the worker and perform emotion analysis. Based on the analysis results, a safety warning is issued immediately if the worker is in danger. Furthermore, the real-time analysis data is stored in cloud storage and can be accessed by managers through a dedicated application.

[1639] Example prompts for generative AI models

[1640] Design a system that uses facial recognition and emotion analysis to monitor the safety of workers in a factory. A high-resolution camera captures images of workers working and analyzes them in real time using an emotion engine. If a worker is in danger or overly fatigued, an alert will be issued and the data will be stored in the cloud. Please provide a concrete code example.

[1641] As a result, a system can be provided that can monitor the emotional state and fatigue level of workers in real time and respond quickly to dangerous situations.

[1642] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1643] Step 1:

[1644] Users install high-resolution cameras in the work area of ​​their factories, which capture detailed images of workers' faces and the work they are doing.

[1645] Input: spatial information (camera installation position), high-resolution camera

[1646] Output: Real-time video data

[1647] Step 2:

[1648] The device transmits the captured real-time video data to a server via the Internet, where streaming technology is used to efficiently transmit large amounts of data.

[1649] Input: Real-time video data

[1650] Output: Video data sent to the server

[1651] Step 3:

[1652] The server uses OpenCV to analyze the received video data and identifies the worker's face using a facial recognition algorithm, which then compares it with pre-registered facial information to identify each worker.

[1653] Input: Transmitted video data, pre-registered face data

[1654] Output: Video data including identified workers (with location information)

[1655] Step 4:

[1656] The server tracks the face of the identified worker and records their position in each frame, allowing the worker to be tracked in real time even if they move within the camera's field of view.

[1657] Input: Video data containing identified workers

[1658] Output: Video data with location information of tracked workers

[1659] Step 5:

[1660] The server automatically generates a clip of the worker's face, zooming and adjusting the frame as needed, using an AI model to generate the optimal clip.

[1661] Input: Location-based video data of tracked workers

[1662] Output: Automatically generated clip video data

[1663] Step 6:

[1664] The server uses an emotion engine to analyze the emotions of the workers in the video clips, specifically identifying emotional states such as smiling or surprised using TensorFlow / Keras.

[1665] Input: Clip video data

[1666] Output: Parsed emotion data (with emotion labels)

[1667] Step 7:

[1668] The server then issues safety alerts based on the analyzed emotion data, for example, visual and audio alerts if a worker is in a dangerous emotional state (anger or sadness).

[1669] Input: Emotion-labeled analysis data

[1670] Output: Safety warning (alert)

[1671] Step 8:

[1672] The server uploads the analyzed data and generated video clips to cloud storage, and also transmits the analyzed data in real time to a remote administrator terminal.

[1673] Input: Analysis data, clip video data

[1674] Output: Data stored in the cloud, analysis data sent to the administrator's terminal

[1675] Step 9:

[1676] Users (administrators) can access cloud storage and check stored data and video clips using a dedicated application. The application has an intuitive interface, allowing users to quickly obtain the information they need.

[1677] Input: Data stored in the cloud (video clips, analysis data)

[1678] Output: Monitoring information available to administrators (interface operations)

[1679] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1680] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1681] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1682] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1683] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1684] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1685] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1686] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1687] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1688] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1689] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1690] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1691] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1692] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1693] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1694] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1695] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1696] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1697] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1698] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1699] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1700] The following is further disclosed regarding the above embodiment.

[1701] (Claim 1)

[1702] means for capturing video of the entire event using a high resolution image capture device;

[1703] means for transmitting the video to a server in real time;

[1704] a means for identifying the subject using a facial recognition algorithm on the server;

[1705] means for tracking the subject within the video;

[1706] means for automatically generating clips focused on the subject;

[1707] means for editing said clips to generate individual video files;

[1708] A means for uploading the generated video file to cloud storage;

[1709] means for providing a user interface that allows access to the video file;

[1710] A system including:

[1711] (Claim 2)

[1712] 10. The system of claim 1, further comprising means for automatically zooming and frame correcting the subject.

[1713] (Claim 3)

[1714] The system of claim 1 further comprising means for pre-registering a facial image of the subject.

[1715]

[1716] "Example 1"

[1717] (Claim 1)

[1718] means for capturing video of the entire event using a high resolution image capture device;

[1719] means for transmitting said video to a data processing device in real time;

[1720] means for identifying the subject using a facial recognition algorithm on a data processing device;

[1721] means for tracking the subject within the video;

[1722] means for automatically generating clips focused on the subject;

[1723] means for editing said clips to generate individual video files;

[1724] A means for uploading the generated video file to cloud storage;

[1725] means for providing a user interface that allows access to the video file;

[1726] A system including:

[1727] (Claim 2)

[1728] 10. The system of claim 1, further comprising means for automatically zooming and frame correcting the subject.

[1729] (Claim 3)

[1730] The system of claim 1 further comprising means for pre-registering a facial image of the subject.

[1731] "Application Example 1"

[1732] (Claim 1)

[1733] means for capturing video of the entire event using a high resolution image capture device;

[1734] means for transmitting the video to a server in real time;

[1735] a means for identifying the subject using a facial recognition algorithm on the server;

[1736] means for tracking the subject within the video;

[1737] means for automatically generating clips focused on the subject;

[1738] means for editing said clips to generate individual video files;

[1739] A means for uploading the generated video file to cloud storage;

[1740] means for providing a user interface that allows access to the video file;

[1741] A means for automatically capturing and recording a child's reactions and moments of enjoyment during virtual shopping;

[1742] A system including:

[1743] (Claim 2)

[1744] 10. The system of claim 1, further comprising means for automatically zooming and frame correcting the subject.

[1745] (Claim 3)

[1746] The system of claim 1 further comprising means for pre-registering a facial image of the subject.

[1747] "Example 2: Combining Emotion Engines"

[1748] (Claim 1)

[1749] means for capturing video of the entire event using a high resolution image capture device;

[1750] means for transmitting the video to a server in real time;

[1751] a means for identifying the subject using a facial recognition algorithm on the server;

[1752] means for tracking the subject within the video;

[1753] means for automatically generating clips focused on the subject;

[1754] means for editing said clips to generate individual video files;

[1755] A means for uploading the generated video file to cloud storage;

[1756] means for providing a user interface that allows access to the video file;

[1757] a means for recognizing emotions in a video using an emotion engine;

[1758] means for editing the video based on the recognized emotions;

[1759] A system including:

[1760] (Claim 2)

[1761] 10. The system of claim 1, further comprising means for automatically zooming and frame correcting the subject.

[1762] (Claim 3)

[1763] The system of claim 1 further comprising means for pre-registering a facial image of the subject.

[1764] "Application example 2 when combining emotion engines"

[1765] (Claim 1)

[1766] means for capturing video of the entire event using a high resolution image capture device;

[1767] means for transmitting the video to a server in real time;

[1768] a means for identifying the subject using a facial recognition algorithm on the server;

[1769] means for tracking the subject within the video;

[1770] means for automatically generating clips focused on the subject;

[1771] means for editing said clips to generate individual video files;

[1772] A means for uploading the generated video file to cloud storage;

[1773] means for providing a user interface that allows access to the video file;

[1774] A means for analyzing emotions in the video in real time using an emotion engine and issuing a safety warning;

[1775] data transmission means for transmitting analysis data in real time;

[1776] A means for storing the data in cloud storage;

[1777] A system including:

[1778] (Claim 2)

[1779] 10. The system of claim 1, further comprising means for automatically zooming and frame correcting the subject.

[1780] (Claim 3)

[1781] The system of claim 1 further comprising means for pre-registering a facial image of the subject. [Explanation of symbols]

[1782] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for capturing video of the entire event using a high resolution image capture device; means for transmitting the video to a server in real time; a means for identifying the subject using a facial recognition algorithm on the server; means for tracking the subject within the video; means for automatically generating clips focused on the subject; means for editing said clips to generate individual video files; A means for uploading the generated video file to cloud storage; means for providing a user interface that allows access to the video file; A system including:

2. The system of claim 1 further comprising means for automatically zooming and framing the subject.

3. The system of claim 1 , further comprising means for pre-registering a facial image of the subject.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A