system

A system analyzes and replaces unauthorized faces and copyrighted music in videos to ensure legal and natural distribution, addressing portrait rights and copyright issues.

JP2026103421APending Publication Date: 2026-06-24SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

AI Technical Summary

Technical Problem

The increasing possibility of portrait rights and copyright infringement due to the presence of others' faces or background music in video distribution, with conventional solutions impairing the naturalness of the video.

Method used

A system that analyzes facial information in videos, compares it with pre-registered permission information, and replaces unauthorized faces with alternative data, while also replacing copyrighted music with copyright-free music, ensuring legal and natural distribution.

Benefits of technology

Enables the distribution of videos without infringing on portrait rights or copyrights, maintaining video naturalness and user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026103421000001_ABST
    Figure 2026103421000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for analyzing video data and identifying detected facial information by matching it with pre-registered permission information, A means for selecting alternative facial data to replace identified facial information and naturally compositing it into video data, A method for analyzing audio data and replacing identified copyright-risk musical portions with copyright-free music data, A means for processing video and audio data in real time and transmitting it to a communication device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, with the spread of video distribution, the possibility of infringement of portrait rights and copyright caused by the faces of others or background music reflected during distribution has been increasing. In conventional countermeasures, it is common to simply mosaic or mute these elements, and there is a problem that the naturalness of the video is impaired. For this reason, there is a need for a technology that can legally and naturally distribute videos containing the faces or music of others.

Means for Solving the Problems

[0005] This invention provides a means for analyzing facial information of other people captured in a video and comparing it with pre-registered permission information. This identifies restricted facial information, selects alternative facial data, and naturally synthesizes it into the video data, enabling distribution without infringing on portrait rights. It also combines this with a means for analyzing audio data and replacing copyrighted music with copyright-free music data, preventing copyright infringement in audio as well. Furthermore, by pre-registering and transmitting user-permitted facial information to the system, it prevents unnecessary substitutions and enables natural video viewing.

[0006] "Video data" refers to video information recorded in digital format, and serves as the basis for analysis and processing.

[0007] "Facial information" refers to information related to a person's face extracted from video data, and this information is used for identification and matching.

[0008] "Permission information" refers to facial information of individuals that users have registered in advance and are permitted to display as is in videos.

[0009] "Alternate face data" refers to data of another face used to replace an unauthorized face that has been detected.

[0010] "Composition" is a technique in video processing that seamlessly combines different video data to reconstruct a single image.

[0011] "Audio data" refers to sound information recorded in digital format, and is the fundamental data for analyzing and processing music and audio.

[0012] "Copyright risk" refers to any situation or element that could potentially lead to copyright infringement if used.

[0013] "Copyright-free music data" refers to music data that can be used without permission from a specific rights holder and is free from rights restrictions. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] The system for implementing this invention involves a server and a terminal working in cooperation. First, the user takes photos of their own face and the faces of authorized acquaintances using the terminal and sends them to the server as an authorized list. This face information is later stored as reference information used during video processing.

[0036] Next, the user uploads a live stream or recorded video to the server. Upon receiving the video data, the server uses AI technology to automatically extract facial information from each frame of the video. This allows the faces of people appearing in the video to be identified in real time and matched against permission information.

[0037] If, after comparing facial information, a face not present in the allowlist is detected, the server accesses the database and selects appropriate alternative facial data randomly or based on user specifications. The selected face is then synthesized using AI to blend naturally into the video, resulting in no visual inconsistencies.

[0038] The server also analyzes the audio data obtained from the video to identify music that may cause copyright issues. If detected, the server seamlessly replaces the relevant portion with royalty-free music data to generate a seamless audio track.

[0039] For example, when a user is live-streaming in the city, if a passerby unexpectedly appears in the frame, their face will be replaced in real time with another face, allowing the stream to continue. Similarly, if background music playing in a cafe or similar location is copyrighted, it will be automatically replaced with different, licensed music. This system performs all processing in the background, allowing users to stream videos with peace of mind without having to go through complicated procedures.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] Users use their devices to take photos of themselves and authorized individuals, registering them as facial recognition data. This data is uploaded from the device to the server as a permission list.

[0043] Step 2:

[0044] Users upload live streams or previously recorded videos to the server. The server stores the received video files for analysis.

[0045] Step 3:

[0046] The server uses AI technology to analyze each frame of the video and detect faces. The detected face data is then compared and matched with the face information in the previously received permission list.

[0047] Step 4:

[0048] The server identifies faces not on the allowed list based on the matching results and selects alternative face data from the database to replace them. The selected alternative face is then seamlessly integrated into the corresponding portion of the video using AI.

[0049] Step 5:

[0050] The server extracts audio data from the video and analyzes the waveform patterns of the music to detect music that poses a copyright risk. If detected, the server selects appropriate copyright-free music from its database.

[0051] Step 6:

[0052] The server replaces the original music portion with detected, royalty-free music data, generating a natural-sounding audio track while maintaining the overall harmony of the audio track.

[0053] Step 7:

[0054] The server provides the processed video data to the user. The user can then securely distribute or download this video.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] Protecting privacy and avoiding copyright issues are crucial challenges in modern video distribution. In particular, the unintended appearance of faces or the use of background music can lead to the leakage of personal information and legal troubles. This invention aims to solve these problems.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes means for analyzing video information and identifying detected face identification information by comparing it with pre-registered permission information; means for selecting alternative face information to replace the identified face identification information and naturally synthesizing it with the video information; and means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information. This makes it possible to protect privacy and avoid copyright issues.

[0060] "Visual information" refers to visual data recorded in the form of images or videos, and this includes image data for each frame.

[0061] "Facial identification information" refers to feature data related to a person's face extracted from video information, and includes feature points such as the position of the eyes, nose, and mouth.

[0062] "Permission information" refers to a list of pre-registered facial identification information, which is used to match facial identification information within video footage.

[0063] "Alternative facial information" refers to artificial or other facial data selected to replace existing facial identification information within video data.

[0064] "Synthesizing" refers to the process of integrating multiple data sets and reconstructing them as a single, natural-looking visual data set.

[0065] "Audio information" refers to audio data accompanying video information, which includes conversations, music, and other sounds.

[0066] "Copyright-free music information" refers to music data that is not subject to copyright restrictions and can be freely used.

[0067] This invention utilizes a system in which a server and a terminal work in cooperation. First, the user acquires facial identification information using the terminal's camera. This information is captured using a dedicated application or camera function on the terminal and transmitted to an information processing device in order to be included in a destination list. The server stores this received facial identification information as permission information in its data storage.

[0068] The user then uploads the live stream or recorded video information to the server. The server processes this video information, using software such as TENSORFLOW® and OpenCV. These tools are used to analyze the video frame by frame and automatically extract face identification information. This allows for the identification of people's faces based on permission information.

[0069] If the identified face is not present in the permission information, the server consults the database and selects alternative face information. The selected face information is then seamlessly integrated into the video information using DeepFake technology or StyleGAN. For example, if a passerby unexpectedly appears in a live stream of a city street, the passerby's face is replaced with another face in real time.

[0070] Furthermore, the server also analyzes audio information to identify music that poses copyright risks. Using AI-powered speech recognition, the music is replaced with royalty-free music. For example, background music playing when filming in a cafe or similar location is automatically replaced with licensed music.

[0071] Here's an example of a specific prompt when using a generative AI model: "I want to develop a system that replaces specific faces in real time with other faces in live-streamed video from a city street. Please tell me what essential technologies are needed. Also, please advise on how to automatically replace specific background music within an audio track."

[0072] This system allows users to protect their privacy, avoid copyright risks, and achieve safe and smooth video streaming.

[0073] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0074] Step 1:

[0075] The user uses the device's camera to acquire facial recognition information. The input is a facial image of the user or an authorized acquaintance. The device uses a facial recognition library such as OpenCV to extract facial feature points. The output is digital data ready to be sent to the server as facial recognition information. Specifically, it provides numerical data representing the positions of the eyes, nose, and mouth.

[0076] Step 2:

[0077] The terminal sends the extracted facial identification information to the server. The input is facial identification information data, and the output is a list of permission information stored in the server's database. The server receives this data using a secure communication protocol (e.g., HTTPS) and stores it in storage to use it as reference information. Specifically, the facial data is encrypted before transmission, and the server verifies data integrity before storing it.

[0078] Step 3:

[0079] The user uploads live stream or recorded video information to the server. The input is a video file captured on the user's device. The server divides this video data frame by frame and uses TensorFlow or OpenCV to extract face identification information from each frame. The output is face identification information for each frame, which is used in the subsequent face matching process. Specifically, it is processed sequentially according to the frame rate, and data is generated in which face regions are identified and extracted.

[0080] Step 4:

[0081] The server compares the extracted face identification information with the permission information to perform identification. The input is the face identification information for each frame, and the output is the result of determining whether it matches or does not match the permission information. A database query is used for the comparison, and as a result, face identification information that does not exist in the permission information is identified. Specifically, the comparison algorithm is used to improve the accuracy of face recognition.

[0082] Step 5:

[0083] If the server detects a face that does not exist in the permission information, it selects alternative face information and composites it into the video. The input is the face identification information that was determined to be a mismatch, and the output is the video with the alternative face composited. DeepFake technology and StyleGAN are used for the synthesis, and the selected alternative face information is integrated into the video in real time. Specifically, the process involves generating different face data and applying it to the video frames in a natural way.

[0084] Step 6:

[0085] The server analyzes the audio information of the video to identify copyright-infringing musical portions. The input is the audio data contained in the video, and the output is audio data with copyright-free music replaced. AI-based music fingerprinting technology is used for the analysis, and the detected musical portions are appropriately replaced. Specifically, the audio track is split, and the problematic portions undergo an overwrite process, including volume adjustment.

[0086] This allows users to distribute videos free from privacy and copyright concerns.

[0087] (Application Example 1)

[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0089] In modern content distribution services, protecting privacy and managing copyrights are critical issues. In particular, the frequent occurrence of unintended individuals appearing on live streams or copyrighted music playing in the background poses legal risks for both streamers and platforms. Addressing these challenges and ensuring compliance while maintaining freedom of distribution is essential.

[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0091] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; means for analyzing audio data and replacing identified copyright-risk music portions with copyright-free music data; and means for processing video and audio data in real time and transmitting it to a communication device. This enables distributors to safely deliver content in real time without concerns about privacy or copyright issues.

[0092] "Video data" refers to visual information acquired from a camera or other recording device, and is composed of multiple frames.

[0093] "Analysis" is the process of breaking down data into smaller parts and revealing its constituent elements and functions.

[0094] "Facial information" refers to data that includes facial feature points and patterns used to identify individuals.

[0095] "Permission information" refers to data and information that has been registered in advance to confirm that specific conditions or criteria are met.

[0096] "Substitution" is the operation of replacing one element with another.

[0097] "Alternative face data" refers to data about another face that is used in place of the original face information, and is designed to be synthesized naturally.

[0098] "Natural synthesis" refers to techniques that ensure that no visual inconsistencies occur when modifications or substitutions are made.

[0099] "Audio data" refers to auditory information obtained from sound acquisition devices such as microphones.

[0100] A "musical portion with copyright risk" refers to a segment of music that is protected by copyright and, if used without permission, could lead to legal problems.

[0101] "Copyright-free music data" refers to music data for which the right to freely use under specific usage conditions has been granted.

[0102] "Real-time" refers to a process where data acquisition, processing, and output are synchronized with the passage of time in the real world.

[0103] A "communication device" is a device or equipment that sends and receives data and has the ability to transmit information over a network.

[0104] The system that realizes this invention operates with a server and a terminal working together. It starts when a user uses a smartphone or other communication terminal to record live video and sends the data to the server. In this process, the terminal acquires video and audio data in real time from the camera and microphone.

[0105] The server analyzes video data using AI technologies such as Amazon Recognition, detects facial information, and compares it with pre-registered permission information. If unauthorized facial information is included, the server selects alternative facial data from the database and uses software such as Unity to seamlessly composite it into the video data. This ensures that even if a specific individual is captured in the video, their face is replaced with another face without any visual incongruity.

[0106] For audio data, AI technology is used to analyze identified musical portions, and if music with copyright risks is included, it is replaced with copyright-free music data. This audio replacement process utilizes speech synthesis services such as Amazon Polly. This avoids legal issues related to copyright when videos are distributed.

[0107] For example, if a user is live-streaming in the middle of town, it doesn't matter if passersby appear in the camera's view or if radio music is playing in the background. Because the server handles all processing in real time, the user can continue streaming without being aware of these issues.

[0108] An example of a prompt sentence to input into the generating AI model is, "Ideas for real-time video editing when a friend accidentally appears in a live stream with your pet."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The user records live video and audio using a communication terminal. Raw video and audio data from the camera and microphone are used as input. This data is streamed to the server in real time.

[0112] Step 2:

[0113] The server processes the received video data using image analysis technologies such as Amazon Recognition. The input is the raw video data sent in step 1. The server analyzes this data and extracts face information. The output is the state in which the face information has been identified.

[0114] Step 3:

[0115] The server compares the detected facial information with existing authorization information. Here, the facial information output in step 2 is used as input. If unauthorized facial information is included, the server selects alternative facial data from a pre-registered database. The output is a pair of the alternative facial data and the identified facial information.

[0116] Step 4:

[0117] The server uses compositing tools such as Unity to seamlessly composite the replacement face data onto the video data. The input for this step is the replacement face data and video data obtained in step 3. These are combined to generate video data that looks natural and doesn't feel unnatural. The output is the replaced video data.

[0118] Step 5:

[0119] The server analyzes the received audio data using AI technology. It analyzes this data using audio analysis technologies such as Amazon Polly. The input is the audio data transferred in step 1. The server detects the music portion that poses copyright risk from this data. The output is the audio data with the identified risk.

[0120] Step 6:

[0121] The server replaces audio data identified as having copyright risks with royalty-free music data. The audio data output in Step 5 and royalty-free music data obtained from the database are used as input. These are combined to generate a seamless audio track. The output is the replaced audio data.

[0122] Step 7:

[0123] The server combines the processed video and audio data and sends it back to the communication device. The input for this step is the final video and audio data after each processing step. The output is the processed live stream sent back to the user terminal, enabling real-time distribution.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] This invention is a system that improves the user experience while avoiding issues related to portrait rights and copyrights by having a server, terminal, and user cooperate to analyze and process video and audio data within video content. In particular, by incorporating an emotion engine, it is possible to dynamically perform appropriate processing according to the context of the video.

[0126] In implementing the system, users first register the facial information they wish to allow via their device. This involves taking a photo of their face using the camera through a simple application operation, and uploading it to the server as part of the allow list. Furthermore, users can either upload videos to the server or configure live streaming settings.

[0127] The server analyzes the received video data and uses AI technology to recognize faces from the video frames. It then compares this face information with an allowed list, and if an unregistered face is found, it selects a random or situationally designated alternative face from the database and seamlessly integrates it into the video.

[0128] In addition, this system is equipped with an emotion engine that analyzes the user's facial expressions and gestures in the video to recognize their emotional state. Based on the recognized emotion, it dynamically changes the video's presentation—for example, music selection or face replacement—synchronizing the video's atmosphere with the emotion to provide a more natural and engaging viewing experience.

[0129] Furthermore, the server also analyzes the audio data in parallel to detect music that involves copyright. It replaces that portion with copyright-free music data stored in the database to generate a seamless audio track.

[0130] For example, if a user is smiling and live-tweeting an event, the emotion engine recognizes this cheerfulness and switches the background music to something bright and energetic. Conversely, if the user's emotions are calm, the music is also changed to a more appropriate, quiet melody. In this way, the system can harmonize the video and audio to match the user's psychological state through emotion recognition.

[0131] In this way, the system based on the present invention utilizes user permission data and emotion recognition to enable comprehensive management that ensures video content is emotionally engaging for viewers while complying with legal and ethical standards.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] Users use their devices to take photos of their own face and the faces of authorized individuals, registering them as a permission list. This data is uploaded from the device to the server.

[0135] Step 2:

[0136] The user either configures a live stream or uploads a recorded video file to the server. The server then begins preparing to analyze the received video data.

[0137] Step 3:

[0138] The server uses AI technology to analyze each frame of the video and detect human faces that appear in the footage. The detected face data is then compared with face information from a pre-submitted permission list.

[0139] Step 4:

[0140] The server identifies unauthorized facial information and selects alternative facial data from the database to replace it. This alternative face is processed using AI technology so that it is seamlessly integrated into the video.

[0141] Step 5:

[0142] The server activates an emotion engine and recognizes the user's emotional state by analyzing their facial expressions and gestures in the video.

[0143] Step 6:

[0144] The server selects and adjusts the music and surrogate face data used in the video based on the recognized emotional state. This synchronizes the mood of the video with the user's emotions.

[0145] Step 7:

[0146] The server simultaneously analyzes the audio data and detects copyrighted music. It then replaces that music with copyright-free music from its database to generate a natural-sounding audio track.

[0147] Step 8:

[0148] The server delivers the processed video to the user. The user can then stream or download this processed video in real time.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0151] Issues related to portrait rights and copyrights that arise when using video content prevent users from creating and sharing content with peace of mind. Furthermore, there is a need to effectively visualize emotional expressions within videos to improve the user experience. Additionally, traditional systems often provide static content that doesn't match the user's emotions.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; and means for analyzing audio data and replacing identified copyright-risky music portions with copyright-free music data. This makes it possible to create video content that is engaging and synchronized with the user's emotions while avoiding issues related to portrait rights and copyright.

[0154] "Analyzing video data" means breaking down each frame of a video and processing it to identify specific elements.

[0155] "Identifying facial information" means extracting facial features detected from video data and comparing them with pre-registered data to check for a match.

[0156] "Selecting alternative face data" means that if an unauthorized face is present in the video, alternative face data will be extracted from the database.

[0157] "Analyzing audio data" refers to the process used to examine the audio track within a video and separate and identify specific music or audio components.

[0158] "Replacing with royalty-free music data" means changing the detected copyrighted music portion to music data that is free from licensing restrictions.

[0159] "Emotion recognition technology" is a technology that identifies a user's emotional state from their facial expressions and actions through image analysis.

[0160] "Adjusting the video's presentation" means dynamically changing elements within the video, such as music and visual effects, in response to the user's emotional state.

[0161] This invention provides a system that effectively avoids issues related to portrait rights and copyrights in video content and improves the user experience through the cooperation of a server, terminal, and user. In particular, by combining image analysis technology and emotion recognition technology and dynamically adjusting the content, it enables output that is more tailored to individual situations.

[0162] First, users register their facial information using their device. This is done using the photo-taking function installed on the device and through a facial recognition application. This image data is uploaded to the server and added to an allow list. Users who wish to process video data either upload the video data to the server via the internet or set up live streaming.

[0163] The server uses AI technology to analyze received video data and has the capability to recognize faces in the video frame by frame. During this process, face information within each frame is checked and compared against an allowed list. If face information is not registered, a substitute face is searched from the database and seamlessly integrated into the video data.

[0164] Emotion recognition technology analyzes users' facial expressions and actions in videos to recognize their emotional state. Based on the recognized emotional information, the server changes the video's presentation, specifically the background music and visual effects, according to the user. This synchronizes the video's atmosphere with the individual user's emotions, making it more engaging for viewers.

[0165] For example, consider a scenario where a user is filming their child's birthday party. The emotion engine recognizes bright smiles in the video and automatically changes the background music to a cheerful one to match the emotion. Similarly, if the scene is emotional, it adjusts the music to a calmer melody.

[0166] The generative AI model is given instructions using prompts. For example, specific instructions such as, "The user is filming their child's birthday party. The scene contains many smiles and celebratory gestures. Select music that suits this scene and appropriately process the unregistered facial information," are possible.

[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0168] Step 1:

[0169] User facial information registration

[0170] The user takes a photo of their face using the device. This input is the captured image data. The device uploads the image to the server via the application and registers it in the allow list. Specifically, the user activates the device's camera and takes a photo of their face following the on-screen instructions. After taking the photo, the application automatically sends the image to the server. The output is the face data registered in the server's allow list.

[0171] Step 2:

[0172] Video upload or live stream settings

[0173] The user selects the video they want to analyze and either uploads it to the server or configures live streaming settings. The input is the video file or streaming information selected by the user. The device then sends this information to the server. Specifically, the user opens a file selection screen on their device, chooses a video, and presses the upload button. The output is the video data stored on the server.

[0174] Step 3:

[0175] Video data analysis and facial recognition

[0176] The server analyzes the received video data frame by frame using AI technology to recognize faces. The input is each frame of the video, and the output is the recognized face information and its location. Specifically, the AI ​​model scans each frame of the video and detects faces. The server then uses this information to prepare for the next step.

[0177] Step 4:

[0178] Facial information matching and surname face synthesis

[0179] The server compares the recognized facial information against an allow list. The input is the output of step 3. If an unregistered face is found as a result of the comparison, the server selects an appropriate substitute face from the database. Specifically, the server then performs a synthesis process and integrates the selected substitute face into the video. At this time, the substitute face is adjusted to match the facial expression and angle of the original video. The output is new video data with corrected facial information.

[0180] Step 5:

[0181] Emotion recognition and performance adjustment

[0182] The server uses emotion recognition technology to analyze the user's facial expressions and gestures in the video and identify their emotional state. The input consists of the original video data and face recognition data. The output is the recognized emotion information, which is used to change the video's music and effects. For example, if the server detects a smile from the user, it will switch to upbeat music that matches the smile. The specific operation involves facial emotion analysis by an AI model and automatic music selection from a music library.

[0183] Step 6:

[0184] Audio data analysis and music replacement

[0185] The server analyzes the audio track and detects music that may infringe on copyright. The input is the audio data of the video. If there is a copyright issue, the affected portion is replaced with copyright-free music. Specifically, the server scans the audio data, detects specific musical phrases, and replaces them. The output is a new video file containing the adjusted audio track.

[0186] (Application Example 2)

[0187] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0188] In the production and distribution of video content, there is a need to mitigate the risks of infringing on portrait rights and copyrights, while also creating emotionally engaging content for viewers. However, current technology lacks the systems to simultaneously satisfy these requirements. Specifically, there are risks such as individuals' faces appearing without permission, the use of copyrighted music, and the content's atmosphere not matching what viewers expect.

[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0190] In this invention, the server includes means for analyzing video information and identifying detected facial information by comparing it with pre-registered permission data; means for selecting alternative facial information to replace the identified facial information and naturally synthesizing it with the video information; means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information; and means for analyzing facial and movement information in the video using an emotion analysis function and dynamically changing the video and audio presentation according to the recognized emotional state. This makes it possible for content creators and distributors to produce and provide emotionally appealing video content to viewers while avoiding issues related to portrait rights and copyrights.

[0191] "Video information" refers to the visual data displayed on the screen, including video content and its individual frames.

[0192] "Facial information" refers to identifiable data about a person's face included in video information, including the position and characteristics of the face.

[0193] "Permission data" refers to information that has been registered in advance and confirmed to be permitted within the system, and specifically refers to facial information for which permission has been granted.

[0194] "Alternative facial information" refers to facial information used to replace facial information that does not match the permission data.

[0195] "Audio information" refers to digital data including voice and music, and specifically to sound information provided in conjunction with video information.

[0196] "Copyright-free music information" refers to music data that is not restricted by copyright and can be used freely.

[0197] "Emotional analysis function" refers to data and algorithms that analyze and infer an individual's emotional state from video and audio information.

[0198] "Dynamic modification" refers to the real-time editing and reconfiguration of video and audio data in response to specific conditions.

[0199] The system for realizing an application of this invention operates through the cooperation of three parties: a server, a terminal, and a user. The user uses the terminal to record video and uploads it to the server. The terminal uses a camera application with facial recognition capabilities to capture facial information authorized by the user and registers it with the server as authorized data in advance.

[0200] The server uses deep learning technology to identify faces in order to analyze the received video information. Specifically, it analyzes the video information frame by frame, compares the detected faces with permission data, selects alternative faces from the database for unregistered faces, and synthesizes them naturally using libraries such as OpenCV. This process makes it possible to avoid issues related to portrait rights.

[0201] Furthermore, the server analyzes the audio information to identify music portions that pose copyright risks. The FFmpeg library is used to replace the detected portions with copyright-free music information. In doing so, AWS (registered trademark) cloud services with user emotion analysis capabilities are used to analyze the user's facial expressions and actions in the video. Emotion analysis technology makes it possible to dynamically change the video and audio effects in real time according to the recognized emotional state.

[0202] For example, if a user uploads a video they filmed while traveling to the server, the server will recognize the user's smile and automatically insert upbeat background music that matches the video. In this way, it is possible to provide emotionally appealing content to viewers.

[0203] An example of a prompt might be, "Please create a video with many smiling user images and add cheerful background music." This prompt can be used as input for a generative AI model.

[0204] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0205] Step 1:

[0206] The user records a video using their device. During this process, the user uses a camera application to record the video and generate video data. The recorded video is immediately saved to the device, and preparations for facial recognition to extract facial information are also initiated within the same application.

[0207] Step 2:

[0208] The user registers the facial information they wish to allow using their device. Specifically, they take a picture of their face using an application with facial recognition capabilities and store it as permission data. The input is facial image data, and the output is a list of permission data sent to the server.

[0209] Step 3:

[0210] The user uploads the captured video information from their device to the server. The device retrieves the video data and sends it to the server via the network. The input is the video data, and the output is the video information stored on the server.

[0211] Step 4:

[0212] The server analyzes the received video information and performs face recognition. It uses a deep learning algorithm to detect face information in each frame. The input is the video data on the server, and the output is a list of detected face information.

[0213] Step 5:

[0214] The server compares the detected face information with the permission data list and generates alternative face information for unregistered faces. If no match is found, information is selected from the alternative face database and synthesized into the video data using OpenCV or similar tools. The output is new video data with the alternative face applied.

[0215] Step 6:

[0216] The server analyzes the audio information and identifies the music portions that pose copyright risks. After identifying the risky portions through audio analysis, it performs replacement processing using the FFmpeg library. The input is audio data on the server, and the output is audio data with royalty-free music set as the input.

[0217] Step 7:

[0218] The server uses emotion analysis capabilities to analyze facial expressions and motion information within the video. It uses services such as AWS Rekognition to estimate the user's emotions. The input is video data on the server, and the output is the estimated emotional state.

[0219] Step 8:

[0220] The server dynamically changes the video and music effects according to the recognized emotional state. For example, if a smile is detected, the background music will be changed to a bright and energetic one. The input is the estimated emotional state and royalty-free music data, and the output is the final video data with the effects changed.

[0221] Through these steps, content creators can avoid legal risks while providing emotionally engaging video content for viewers.

[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0225] [Second Embodiment]

[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0238] The system for implementing this invention involves a server and a terminal working in cooperation. First, the user takes photos of their own face and the faces of authorized acquaintances using the terminal and sends them to the server as an authorized list. This face information is later stored as reference information used during video processing.

[0239] Next, the user uploads a live stream or recorded video to the server. Upon receiving the video data, the server uses AI technology to automatically extract facial information from each frame of the video. This allows the faces of people appearing in the video to be identified in real time and matched against permission information.

[0240] If, after comparing facial information, a face not present in the allowlist is detected, the server accesses the database and selects appropriate alternative facial data randomly or based on user specifications. The selected face is then synthesized using AI to blend naturally into the video, resulting in no visual inconsistencies.

[0241] The server also analyzes the audio data obtained from the video to identify music that may cause copyright issues. If detected, the server seamlessly replaces the relevant portion with royalty-free music data to generate a seamless audio track.

[0242] For example, when a user is live-streaming in the city, if a passerby unexpectedly appears in the frame, their face will be replaced in real time with another face, allowing the stream to continue. Similarly, if background music playing in a cafe or similar location is copyrighted, it will be automatically replaced with different, licensed music. This system performs all processing in the background, allowing users to stream videos with peace of mind without having to go through complicated procedures.

[0243] The following describes the processing flow.

[0244] Step 1:

[0245] Users use their devices to take photos of themselves and authorized individuals, registering them as facial recognition data. This data is uploaded from the device to the server as a permission list.

[0246] Step 2:

[0247] Users upload live streams or previously recorded videos to the server. The server stores the received video files for analysis.

[0248] Step 3:

[0249] The server uses AI technology to analyze each frame of the video and detect faces. The detected face data is then compared and matched with the face information in the previously received permission list.

[0250] Step 4:

[0251] The server identifies faces not on the allowed list based on the matching results and selects alternative face data from the database to replace them. The selected alternative face is then seamlessly integrated into the corresponding portion of the video using AI.

[0252] Step 5:

[0253] The server extracts audio data from the video and analyzes the waveform patterns of the music to detect music that poses a copyright risk. If detected, the server selects appropriate copyright-free music from its database.

[0254] Step 6:

[0255] The server replaces the original music portion with detected, royalty-free music data, generating a natural-sounding audio track while maintaining the overall harmony of the audio track.

[0256] Step 7:

[0257] The server provides the processed video data to the user. The user can then securely distribute or download this video.

[0258] (Example 1)

[0259] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0260] Protecting privacy and avoiding copyright issues are crucial challenges in modern video distribution. In particular, the unintended appearance of faces or the use of background music can lead to the leakage of personal information and legal troubles. This invention aims to solve these problems.

[0261] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0262] In this invention, the server includes means for analyzing video information and identifying detected face identification information by comparing it with pre-registered permission information; means for selecting alternative face information to replace the identified face identification information and naturally synthesizing it with the video information; and means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information. This makes it possible to protect privacy and avoid copyright issues.

[0263] "Visual information" refers to visual data recorded in the form of images or videos, and this includes image data for each frame.

[0264] "Facial identification information" refers to feature data related to a person's face extracted from video information, and includes feature points such as the position of the eyes, nose, and mouth.

[0265] "Permission information" refers to a list of pre-registered facial identification information, which is used to match facial identification information within video footage.

[0266] "Alternative facial information" refers to artificial or other facial data selected to replace existing facial identification information within video data.

[0267] "Synthesizing" refers to the process of integrating multiple data sets and reconstructing them as a single, natural-looking visual data set.

[0268] "Audio information" refers to audio data accompanying video information, which includes conversations, music, and other sounds.

[0269] "Copyright-free music information" refers to music data that is not subject to copyright restrictions and can be freely used.

[0270] This invention utilizes a system in which a server and a terminal work in cooperation. First, the user acquires facial identification information using the terminal's camera. This information is captured using a dedicated application or camera function on the terminal and transmitted to an information processing device in order to be included in a destination list. The server stores this received facial identification information as permission information in its data storage.

[0271] The user then uploads the live stream or recorded video information to the server. The server processes this video information, using software such as TensorFlow and OpenCV. These tools are used to analyze the video frame by frame and automatically extract face identification information. This allows for the identification of people's faces based on permission information.

[0272] If the identified face is not present in the permission information, the server consults the database and selects alternative face information. The selected face information is then seamlessly integrated into the video information using DeepFake technology or StyleGAN. For example, if a passerby unexpectedly appears in a live stream of a city street, the passerby's face is replaced with another face in real time.

[0273] Furthermore, the server also analyzes audio information to identify music that poses copyright risks. Using AI-powered speech recognition, the music is replaced with royalty-free music. For example, background music playing when filming in a cafe or similar location is automatically replaced with licensed music.

[0274] Here's an example of a specific prompt when using a generative AI model: "I want to develop a system that replaces specific faces in real time with other faces in live-streamed video from a city street. Please tell me what essential technologies are needed. Also, please advise on how to automatically replace specific background music within an audio track."

[0275] This system allows users to protect their privacy, avoid copyright risks, and achieve safe and smooth video streaming.

[0276] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0277] Step 1:

[0278] The user uses the device's camera to acquire facial recognition information. The input is a facial image of the user or an authorized acquaintance. The device uses a facial recognition library such as OpenCV to extract facial feature points. The output is digital data ready to be sent to the server as facial recognition information. Specifically, it provides numerical data representing the positions of the eyes, nose, and mouth.

[0279] Step 2:

[0280] The terminal sends the extracted facial identification information to the server. The input is facial identification information data, and the output is a list of permission information stored in the server's database. The server receives this data using a secure communication protocol (e.g., HTTPS) and stores it in storage to use it as reference information. Specifically, the facial data is encrypted before transmission, and the server verifies data integrity before storing it.

[0281] Step 3:

[0282] The user uploads live stream or recorded video information to the server. As input, there is a video file captured by the user's terminal. The server splits this video data frame by frame and uses TensorFlow or OpenCV to extract face identification information from each frame. The output is the face identification information for each frame, which is used in the subsequent face matching process. Specifically, it is processed sequentially according to the frame rate, and data identifying and extracting the face area is generated.

[0283] Step 4:

[0284] The server matches the extracted face identification information with the permission information for identification. The input is the face identification information for each frame, and the output is the determination result of match or mismatch with the permission information. A database query is used for the matching, and as a result, face identification information that does not exist in the permission information is identified. As a specific operation, a matching algorithm is utilized to improve face recognition accuracy.

[0285] Step 5:

[0286] If the server discovers a face that does not exist in the permission information, it selects alternative face information and synthesizes it into the video. The input is the face identification information determined to be a mismatch, and the output is the video with the alternative face synthesized. DeepFake technology or StyleGAN is used for the synthesis, and the selected alternative face information is integrated into the video in real time. Specifically, different face data is generated and processed to be naturally applied to the video frames.

[0287] Step 6:

[0288] The server analyzes the audio information of the video to identify music parts with copyright risks. The input is the audio data included in the video, and the output is the audio data with copyright-free music replaced. Music fingerprint recognition technology using AI is utilized for the analysis, and the detected music parts are appropriately replaced. As a specific operation, the audio track is split, and the problematic parts undergo an overwriting process including volume adjustment.

[0289] This allows users to distribute videos free from privacy and copyright concerns.

[0290] (Application Example 1)

[0291] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0292] In modern content distribution services, protecting privacy and managing copyrights are critical issues. In particular, the frequent occurrence of unintended individuals appearing on live streams or copyrighted music playing in the background poses legal risks for both streamers and platforms. Addressing these challenges and ensuring compliance while maintaining freedom of distribution is essential.

[0293] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0294] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; means for analyzing audio data and replacing identified copyright-risk music portions with copyright-free music data; and means for processing video and audio data in real time and transmitting it to a communication device. This enables distributors to safely deliver content in real time without concerns about privacy or copyright issues.

[0295] "Video data" refers to visual information acquired from a camera or other recording device, and is composed of multiple frames.

[0296] "Analysis" is the process of breaking down data into smaller parts and revealing its constituent elements and functions.

[0297] "Facial information" refers to data that includes facial feature points and patterns used to identify individuals.

[0298] "Permission information" refers to data and information that has been registered in advance to confirm that specific conditions or criteria are met.

[0299] "Substitution" is the operation of replacing one element with another.

[0300] "Alternative face data" refers to data about another face that is used in place of the original face information, and is designed to be synthesized naturally.

[0301] "Natural synthesis" refers to techniques that ensure that no visual inconsistencies occur when modifications or substitutions are made.

[0302] "Audio data" refers to auditory information obtained from sound acquisition devices such as microphones.

[0303] "Music portions with copyright risks" refer to segments of music that are protected by copyright and whose use without permission may cause legal problems.

[0304] "Copyright-free music data" refers to music data for which the right to freely use under specific usage conditions has been granted.

[0305] "Real-time" refers to a process where data acquisition, processing, and output are synchronized with the passage of time in the real world.

[0306] A "communication device" is a device or equipment that sends and receives data and has the ability to transmit information over a network.

[0307] The system that realizes this invention operates with the cooperation of a server and a terminal. It is initiated when a user uses a smartphone or other communication terminal to shoot a live video and transmits the data to the server. At this time, the terminal acquires video data and audio data in real time from a camera and a microphone.

[0308] The server analyzes the video data using AI technologies such as Amazon Recognition, detects face information, and compares it with pre-registered permission information. If the face information contains unpermitted content, the server selects alternative face data from the database and synthetically integrates it into the video data naturally using software such as Unity. As a result, even if a specific individual is captured, it is replaced with another face without a visual sense of incongruity.

[0309] Regarding the audio data, the music parts identified using AI technologies are analyzed, and if the music contains copyright risks, it is replaced with copyright-free music data. For this audio replacement process, voice synthesis services such as Amazon Polly are utilized. As a result, legal issues regarding copyright are avoided when the video is distributed.

[0310] As a specific example, even if a passerby is captured by the camera or radio music is playing in the background while the user is conducting a live stream in the middle of the town, there is no problem. Since the server performs all the processing in real time, the user can continue the distribution without being aware of these problems.

[0311] An example of a prompt sentence input to the generative AI model is "Ideas for real-time video editing when a friend is captured during a live stream with a pet."

[0312] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0313] Step 1:

[0314] The user records live video and audio using a communication terminal. Raw video and audio data from the camera and microphone are used as input. This data is streamed to the server in real time.

[0315] Step 2:

[0316] The server processes the received video data using image analysis technologies such as Amazon Recognition. The input is the raw video data sent in step 1. The server analyzes this data and extracts face information. The output is the state in which the face information has been identified.

[0317] Step 3:

[0318] The server compares the detected facial information with existing authorization information. Here, the facial information output in step 2 is used as input. If unauthorized facial information is included, the server selects alternative facial data from a pre-registered database. The output is a pair of the alternative facial data and the identified facial information.

[0319] Step 4:

[0320] The server uses compositing tools such as Unity to seamlessly composite the replacement face data onto the video data. The input for this step is the replacement face data and video data obtained in step 3. These are combined to generate video data that looks natural and doesn't feel unnatural. The output is the replaced video data.

[0321] Step 5:

[0322] The server analyzes the received audio data using AI technology. It analyzes this data using audio analysis technologies such as Amazon Polly. The input is the audio data transferred in step 1. The server detects the music portion that poses copyright risk from this data. The output is the audio data with the identified risk.

[0323] Step 6:

[0324] The server replaces audio data identified as having copyright risks with royalty-free music data. The audio data output in Step 5 and royalty-free music data obtained from the database are used as input. These are combined to generate a seamless audio track. The output is the replaced audio data.

[0325] Step 7:

[0326] The server combines the processed video and audio data and sends it back to the communication device. The input for this step is the final video and audio data after each processing step. The output is the processed live stream sent back to the user terminal, enabling real-time distribution.

[0327] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0328] This invention is a system that improves the user experience while avoiding issues related to portrait rights and copyrights by having a server, terminal, and user cooperate to analyze and process video and audio data within video content. In particular, by incorporating an emotion engine, it is possible to dynamically perform appropriate processing according to the context of the video.

[0329] In implementing the system, users first register the facial information they wish to allow via their device. This involves taking a photo of their face using the camera through a simple application operation, and uploading it to the server as part of the allow list. Furthermore, users can either upload videos to the server or configure live streaming settings.

[0330] The server analyzes the received video data and uses AI technology to recognize faces from the video frames. It then compares this face information with an allowed list, and if an unregistered face is found, it selects a random or situationally designated alternative face from the database and seamlessly integrates it into the video.

[0331] In addition, this system is equipped with an emotion engine that analyzes the user's facial expressions and gestures in the video to recognize their emotional state. Based on the recognized emotion, it dynamically changes the video's presentation—for example, music selection or face replacement—synchronizing the video's atmosphere with the emotion to provide a more natural and engaging viewing experience.

[0332] Furthermore, the server also analyzes the audio data in parallel to detect music that involves copyright. It replaces that portion with copyright-free music data stored in the database to generate a seamless audio track.

[0333] For example, if a user is smiling and live-tweeting an event, the emotion engine recognizes this cheerfulness and switches the background music to something bright and energetic. Conversely, if the user's emotions are calm, the music is also changed to a more appropriate, quiet melody. In this way, the system can harmonize the video and audio to match the user's psychological state through emotion recognition.

[0334] In this way, the system based on the present invention utilizes user permission data and emotion recognition to enable comprehensive management that ensures video content is emotionally engaging for viewers while complying with legal and ethical standards.

[0335] The following describes the processing flow.

[0336] Step 1:

[0337] Users use their devices to take photos of their own face and the faces of authorized individuals, registering them as a permission list. This data is uploaded from the device to the server.

[0338] Step 2:

[0339] The user either configures a live stream or uploads a recorded video file to the server. The server then begins preparing to analyze the received video data.

[0340] Step 3:

[0341] The server uses AI technology to analyze each frame of the video and detect human faces that appear in the footage. The detected face data is then compared with face information from a pre-submitted permission list.

[0342] Step 4:

[0343] The server identifies unauthorized facial information and selects alternative facial data from the database to replace it. This alternative face is processed using AI technology so that it is seamlessly integrated into the video.

[0344] Step 5:

[0345] The server activates an emotion engine and recognizes the user's emotional state by analyzing their facial expressions and gestures in the video.

[0346] Step 6:

[0347] The server selects and adjusts the music and surrogate face data used in the video based on the recognized emotional state. This synchronizes the mood of the video with the user's emotions.

[0348] Step 7:

[0349] The server simultaneously analyzes the audio data and detects copyrighted music. It then replaces that music with copyright-free music from its database to generate a natural-sounding audio track.

[0350] Step 8:

[0351] The server delivers the processed video to the user. The user can then stream or download this processed video in real time.

[0352] (Example 2)

[0353] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0354] Issues related to portrait rights and copyrights that arise when using video content prevent users from creating and sharing content with peace of mind. Furthermore, there is a need to effectively visualize emotional expressions within videos to improve the user experience. Additionally, traditional systems often provide static content that doesn't match the user's emotions.

[0355] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0356] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; and means for analyzing audio data and replacing identified copyright-risky music portions with copyright-free music data. This makes it possible to create video content that is engaging and synchronized with the user's emotions while avoiding issues related to portrait rights and copyright.

[0357] "Analyzing video data" means breaking down each frame of a video and processing it to identify specific elements.

[0358] "Identifying facial information" means extracting facial features detected from video data and comparing them with pre-registered data to check for a match.

[0359] "Selecting alternative face data" means that if an unauthorized face is present in the video, alternative face data will be extracted from the database.

[0360] "Analyzing audio data" refers to the process used to examine the audio track within a video and separate and identify specific music or audio components.

[0361] "Replacing with royalty-free music data" means changing the detected copyrighted music portion to music data that is free from licensing restrictions.

[0362] "Emotion recognition technology" is a technology that identifies a user's emotional state from their facial expressions and actions through image analysis.

[0363] "Adjusting the video's presentation" means dynamically changing elements within the video, such as music and visual effects, in response to the user's emotional state.

[0364] This invention provides a system that effectively avoids issues related to portrait rights and copyrights in video content and improves the user experience through the cooperation of a server, terminal, and user. In particular, by combining image analysis technology and emotion recognition technology and dynamically adjusting the content, it enables output that is more tailored to individual situations.

[0365] First, users register their facial information using their device. This is done using the photo-taking function installed on the device and through a facial recognition application. This image data is uploaded to the server and added to an allow list. Users who wish to process video data either upload the video data to the server via the internet or set up live streaming.

[0366] The server uses AI technology to analyze received video data and has the capability to recognize faces in the video frame by frame. During this process, face information within each frame is checked and compared against an allowed list. If face information is not registered, a substitute face is searched from the database and seamlessly integrated into the video data.

[0367] Emotion recognition technology analyzes users' facial expressions and actions in videos to recognize their emotional state. Based on the recognized emotional information, the server changes the video's presentation, specifically the background music and visual effects, according to the user. This synchronizes the video's atmosphere with the individual user's emotions, making it more engaging for viewers.

[0368] For example, consider a scenario where a user is filming their child's birthday party. The emotion engine recognizes bright smiles in the video and automatically changes the background music to a cheerful one to match the emotion. Similarly, if the scene is emotional, it adjusts the music to a calmer melody.

[0369] The generative AI model is given instructions using prompts. For example, specific instructions such as, "The user is filming their child's birthday party. The scene contains many smiles and celebratory gestures. Select music that suits this scene and appropriately process the unregistered facial information," are possible.

[0370] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0371] Step 1:

[0372] User facial information registration

[0373] The user takes a photo of their face using the device. This input is the captured image data. The device uploads the image to the server via the application and registers it in the allow list. Specifically, the user activates the device's camera and takes a photo of their face following the on-screen instructions. After taking the photo, the application automatically sends the image to the server. The output is the face data registered in the server's allow list.

[0374] Step 2:

[0375] Video upload or live stream settings

[0376] The user selects the video they want to analyze and either uploads it to the server or configures live streaming settings. The input is the video file or streaming information selected by the user. The device then sends this information to the server. Specifically, the user opens a file selection screen on their device, chooses a video, and presses the upload button. The output is the video data stored on the server.

[0377] Step 3:

[0378] Video data analysis and facial recognition

[0379] The server analyzes the received video data frame by frame using AI technology to recognize faces. The input is each frame of the video, and the output is the recognized face information and its location. Specifically, the AI ​​model scans each frame of the video and detects faces. The server then uses this information to prepare for the next step.

[0380] Step 4:

[0381] Facial information matching and surname face synthesis

[0382] The server compares the recognized facial information against an allow list. The input is the output of step 3. If an unregistered face is found as a result of the comparison, the server selects an appropriate substitute face from the database. Specifically, the server then performs a synthesis process and integrates the selected substitute face into the video. At this time, the substitute face is adjusted to match the facial expression and angle of the original video. The output is new video data with corrected facial information.

[0383] Step 5:

[0384] Emotion recognition and performance adjustment

[0385] The server uses emotion recognition technology to analyze the user's facial expressions and gestures in the video and identify their emotional state. The input consists of the original video data and face recognition data. The output is the recognized emotion information, which is used to change the video's music and effects. For example, if the server detects a smile from the user, it will switch to upbeat music that matches the smile. The specific operation involves facial emotion analysis by an AI model and automatic music selection from a music library.

[0386] Step 6:

[0387] Audio data analysis and music replacement

[0388] The server analyzes the audio track and detects music that may infringe on copyright. The input is the audio data of the video. If there is a copyright issue, the affected portion is replaced with copyright-free music. Specifically, the server scans the audio data, detects specific musical phrases, and replaces them. The output is a new video file containing the adjusted audio track.

[0389] (Application Example 2)

[0390] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0391] In the production and distribution of video content, there is a need to mitigate the risks of infringing on portrait rights and copyrights, while also creating emotionally engaging content for viewers. However, current technology lacks the systems to simultaneously satisfy these requirements. Specifically, there are risks such as individuals' faces appearing without permission, the use of copyrighted music, and the content's atmosphere not matching what viewers expect.

[0392] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0393] In this invention, the server includes means for analyzing video information and identifying detected facial information by comparing it with pre-registered permission data; means for selecting alternative facial information to replace the identified facial information and naturally synthesizing it with the video information; means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information; and means for analyzing facial and movement information in the video using an emotion analysis function and dynamically changing the video and audio presentation according to the recognized emotional state. This makes it possible for content creators and distributors to produce and provide emotionally appealing video content to viewers while avoiding issues related to portrait rights and copyrights.

[0394] "Video information" refers to the visual data displayed on the screen, including video content and its individual frames.

[0395] "Facial information" refers to identifiable data about a person's face included in video information, including the position and characteristics of the face.

[0396] "Permission data" refers to information that has been registered in advance and confirmed to be permitted within the system, and specifically refers to facial information for which permission has been granted.

[0397] "Alternative facial information" refers to facial information used to replace facial information that does not match the permission data.

[0398] "Audio information" refers to digital data including voice and music, and specifically to sound information provided in conjunction with video information.

[0399] "Copyright-free music information" refers to music data that is not restricted by copyright and can be used freely.

[0400] "Emotional analysis function" refers to data and algorithms that analyze and infer an individual's emotional state from video and audio information.

[0401] "Dynamic modification" refers to the real-time editing and reconfiguration of video and audio data in response to specific conditions.

[0402] The system for realizing an application of this invention operates through the cooperation of three parties: a server, a terminal, and a user. The user uses the terminal to record video and uploads it to the server. The terminal uses a camera application with facial recognition capabilities to capture facial information authorized by the user and registers it with the server as authorized data in advance.

[0403] The server uses deep learning technology to identify faces in order to analyze the received video information. Specifically, it analyzes the video information frame by frame, compares the detected faces with permission data, selects alternative faces from the database for unregistered faces, and synthesizes them naturally using libraries such as OpenCV. This process makes it possible to avoid issues related to portrait rights.

[0404] Furthermore, the server analyzes the audio information to identify copyright-infringing musical portions. The FFmpeg library is used to replace these detected portions with copyright-free music. During this process, AWS cloud services with user emotion analysis capabilities are utilized to analyze the user's facial expressions and actions within the video. This emotion analysis technology enables the video and audio to be dynamically modified in real time according to the recognized emotional state.

[0405] For example, if a user uploads a video they filmed while traveling to the server, the server will recognize the user's smile and automatically insert upbeat background music that matches the video. In this way, it is possible to provide emotionally appealing content to viewers.

[0406] An example of a prompt might be, "Please create a video with many smiling user images and add cheerful background music." This prompt can be used as input for a generative AI model.

[0407] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0408] Step 1:

[0409] The user records a video using their device. During this process, the user uses a camera application to record the video and generate video data. The recorded video is immediately saved to the device, and preparations for facial recognition to extract facial information are also initiated within the same application.

[0410] Step 2:

[0411] The user registers the facial information they wish to allow using their device. Specifically, they take a picture of their face using an application with facial recognition capabilities and store it as permission data. The input is facial image data, and the output is a list of permission data sent to the server.

[0412] Step 3:

[0413] The user uploads the captured video information from their device to the server. The device retrieves the video data and sends it to the server via the network. The input is the video data, and the output is the video information stored on the server.

[0414] Step 4:

[0415] The server analyzes the received video information and performs face recognition. It uses a deep learning algorithm to detect face information in each frame. The input is the video data on the server, and the output is a list of detected face information.

[0416] Step 5:

[0417] The server compares the detected face information with the permission data list and generates alternative face information for unregistered faces. If no match is found, information is selected from the alternative face database and synthesized into the video data using OpenCV or similar tools. The output is new video data with the alternative face applied.

[0418] Step 6:

[0419] The server analyzes the audio information and identifies the music portions that pose copyright risks. After identifying the risky portions through audio analysis, it performs replacement processing using the FFmpeg library. The input is audio data on the server, and the output is audio data with royalty-free music set as the input.

[0420] Step 7:

[0421] The server uses emotion analysis capabilities to analyze facial expressions and motion information within the video. It uses services such as AWS Rekognition to estimate the user's emotions. The input is video data on the server, and the output is the estimated emotional state.

[0422] Step 8:

[0423] The server dynamically changes the video and music effects according to the recognized emotional state. For example, if a smile is detected, the background music will be changed to a bright and energetic one. The input is the estimated emotional state and royalty-free music data, and the output is the final video data with the effects changed.

[0424] Through these steps, content creators can avoid legal risks while providing emotionally engaging video content for viewers.

[0425] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0426] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0427] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0428] [Third Embodiment]

[0429] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0430] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0431] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0432] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0433] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0435] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0436] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0437] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0438] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0439] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0440] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0441] The system for implementing this invention involves a server and a terminal working in cooperation. First, the user takes photos of their own face and the faces of authorized acquaintances using the terminal and sends them to the server as an authorized list. This face information is later stored as reference information used during video processing.

[0442] Next, the user uploads a live stream or recorded video to the server. Upon receiving the video data, the server uses AI technology to automatically extract facial information from each frame of the video. This allows the faces of people appearing in the video to be identified in real time and matched against permission information.

[0443] If, after comparing facial information, a face not present in the allowlist is detected, the server accesses the database and selects appropriate alternative facial data randomly or based on user specifications. The selected face is then synthesized using AI to blend naturally into the video, resulting in no visual inconsistencies.

[0444] The server also analyzes the audio data obtained from the video to identify music that may cause copyright issues. If detected, the server seamlessly replaces the relevant portion with royalty-free music data to generate a seamless audio track.

[0445] For example, when a user is live-streaming in the city, if a passerby unexpectedly appears in the frame, their face will be replaced in real time with another face, allowing the stream to continue. Similarly, if background music playing in a cafe or similar location is copyrighted, it will be automatically replaced with different, licensed music. This system performs all processing in the background, allowing users to stream videos with peace of mind without having to go through complicated procedures.

[0446] The following describes the processing flow.

[0447] Step 1:

[0448] Users use their devices to take photos of themselves and authorized individuals, registering them as facial recognition data. This data is uploaded from the device to the server as a permission list.

[0449] Step 2:

[0450] Users upload live streams or previously recorded videos to the server. The server stores the received video files for analysis.

[0451] Step 3:

[0452] The server uses AI technology to analyze each frame of the video and detect faces. The detected face data is then compared and matched with the face information in the previously received permission list.

[0453] Step 4:

[0454] The server identifies faces not on the allowed list based on the matching results and selects alternative face data from the database to replace them. The selected alternative face is then seamlessly integrated into the corresponding portion of the video using AI.

[0455] Step 5:

[0456] The server extracts audio data from the video and analyzes the waveform patterns of the music to detect music that poses a copyright risk. If detected, the server selects appropriate copyright-free music from its database.

[0457] Step 6:

[0458] The server replaces the original music portion with detected, royalty-free music data, generating a natural-sounding audio track while maintaining the overall harmony of the audio track.

[0459] Step 7:

[0460] The server provides the processed video data to the user. The user can then securely distribute or download this video.

[0461] (Example 1)

[0462] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0463] Protecting privacy and avoiding copyright issues are crucial challenges in modern video distribution. In particular, the unintended appearance of faces or the use of background music can lead to the leakage of personal information and legal troubles. This invention aims to solve these problems.

[0464] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0465] In this invention, the server includes means for analyzing video information and identifying detected face identification information by comparing it with pre-registered permission information; means for selecting alternative face information to replace the identified face identification information and naturally synthesizing it with the video information; and means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information. This makes it possible to protect privacy and avoid copyright issues.

[0466] "Visual information" refers to visual data recorded in the form of images or videos, and this includes image data for each frame.

[0467] "Facial identification information" refers to feature data related to a person's face extracted from video information, and includes feature points such as the position of the eyes, nose, and mouth.

[0468] "Permission information" refers to a list of pre-registered facial identification information, which is used to match facial identification information within video footage.

[0469] "Alternative facial information" refers to artificial or other facial data selected to replace existing facial identification information within video data.

[0470] "Synthesizing" refers to the process of integrating multiple data sets and reconstructing them as a single, natural-looking visual data set.

[0471] "Audio information" refers to audio data accompanying video information, which includes conversations, music, and other sounds.

[0472] "Copyright-free music information" refers to music data that is not subject to copyright restrictions and can be freely used.

[0473] This invention utilizes a system in which a server and a terminal work in cooperation. First, the user acquires facial identification information using the terminal's camera. This information is captured using a dedicated application or camera function on the terminal and transmitted to an information processing device in order to be included in a destination list. The server stores this received facial identification information as permission information in its data storage.

[0474] The user then uploads the live stream or recorded video information to the server. The server processes this video information, using software such as TensorFlow and OpenCV. These tools are used to analyze the video frame by frame and automatically extract face identification information. This allows for the identification of people's faces based on permission information.

[0475] If the identified face is not present in the permission information, the server consults the database and selects alternative face information. The selected face information is then seamlessly integrated into the video information using DeepFake technology or StyleGAN. For example, if a passerby unexpectedly appears in a live stream of a city street, the passerby's face is replaced with another face in real time.

[0476] Furthermore, the server also analyzes audio information to identify music that poses copyright risks. Using AI-powered speech recognition, the music is replaced with royalty-free music. For example, background music playing when filming in a cafe or similar location is automatically replaced with licensed music.

[0477] Here's an example of a specific prompt when using a generative AI model: "I want to develop a system that replaces specific faces in real time with other faces in live-streamed video from a city street. Please tell me what essential technologies are needed. Also, please advise on how to automatically replace specific background music within an audio track."

[0478] This system allows users to protect their privacy, avoid copyright risks, and achieve safe and smooth video streaming.

[0479] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0480] Step 1:

[0481] The user uses the device's camera to acquire facial recognition information. The input is a facial image of the user or an authorized acquaintance. The device uses a facial recognition library such as OpenCV to extract facial feature points. The output is digital data ready to be sent to the server as facial recognition information. Specifically, it provides numerical data representing the positions of the eyes, nose, and mouth.

[0482] Step 2:

[0483] The terminal sends the extracted facial identification information to the server. The input is facial identification information data, and the output is a list of permission information stored in the server's database. The server receives this data using a secure communication protocol (e.g., HTTPS) and stores it in storage to use it as reference information. Specifically, the facial data is encrypted before transmission, and the server verifies data integrity before storing it.

[0484] Step 3:

[0485] The user uploads live stream or recorded video information to the server. The input is a video file captured on the user's device. The server divides this video data frame by frame and uses TensorFlow or OpenCV to extract face identification information from each frame. The output is face identification information for each frame, which is used in the subsequent face matching process. Specifically, it is processed sequentially according to the frame rate, and data is generated in which face regions are identified and extracted.

[0486] Step 4:

[0487] The server compares the extracted face identification information with the permission information to perform identification. The input is the face identification information for each frame, and the output is the result of determining whether it matches or does not match the permission information. A database query is used for the comparison, and as a result, face identification information that does not exist in the permission information is identified. Specifically, the comparison algorithm is used to improve the accuracy of face recognition.

[0488] Step 5:

[0489] If the server detects a face that does not exist in the permission information, it selects alternative face information and composites it into the video. The input is the face identification information that was determined to be a mismatch, and the output is the video with the alternative face composited. DeepFake technology and StyleGAN are used for the synthesis, and the selected alternative face information is integrated into the video in real time. Specifically, the process involves generating different face data and applying it to the video frames in a natural way.

[0490] Step 6:

[0491] The server analyzes the audio information of the video to identify copyright-infringing musical portions. The input is the audio data contained in the video, and the output is audio data with copyright-free music replaced. AI-based music fingerprinting technology is used for the analysis, and the detected musical portions are appropriately replaced. Specifically, the audio track is split, and the problematic portions undergo an overwrite process, including volume adjustment.

[0492] This allows users to distribute videos free from privacy and copyright concerns.

[0493] (Application Example 1)

[0494] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0495] In modern content distribution services, protecting privacy and managing copyrights are critical issues. In particular, the frequent occurrence of unintended individuals appearing on live streams or copyrighted music playing in the background poses legal risks for both streamers and platforms. Addressing these challenges and ensuring compliance while maintaining freedom of distribution is essential.

[0496] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0497] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; means for analyzing audio data and replacing identified copyright-risk music portions with copyright-free music data; and means for processing video and audio data in real time and transmitting it to a communication device. This enables distributors to safely deliver content in real time without concerns about privacy or copyright issues.

[0498] "Video data" refers to visual information acquired from a camera or other recording device, and is composed of multiple frames.

[0499] "Analysis" is the process of breaking down data into smaller parts and revealing its constituent elements and functions.

[0500] "Facial information" refers to data that includes facial feature points and patterns used to identify individuals.

[0501] "Permission information" refers to data and information that has been registered in advance to confirm that specific conditions or criteria are met.

[0502] "Substitution" is the operation of replacing one element with another.

[0503] "Alternative face data" refers to data about another face that is used in place of the original face information, and is designed to be synthesized naturally.

[0504] "Natural synthesis" refers to techniques that ensure that no visual inconsistencies occur when modifications or substitutions are made.

[0505] "Audio data" refers to auditory information obtained from sound acquisition devices such as microphones.

[0506] "Music portions with copyright risks" refer to segments of music that are protected by copyright and whose use without permission may cause legal problems.

[0507] "Copyright-free music data" refers to music data for which the right to freely use under specific usage conditions has been granted.

[0508] "Real-time" refers to a process where data acquisition, processing, and output are synchronized with the passage of time in the real world.

[0509] A "communication device" is a device or equipment that sends and receives data and has the ability to transmit information over a network.

[0510] The system that realizes this invention operates with a server and a terminal working together. It starts when a user uses a smartphone or other communication terminal to record live video and sends the data to the server. In this process, the terminal acquires video and audio data in real time from the camera and microphone.

[0511] The server analyzes video data using AI technologies such as Amazon Recognition, detects facial information, and compares it with pre-registered permission information. If unauthorized facial information is included, the server selects alternative facial data from the database and uses software such as Unity to seamlessly composite it into the video data. This ensures that even if a specific individual is captured in the video, their face is replaced with another face without any visual incongruity.

[0512] For audio data, AI technology is used to analyze identified musical portions, and if music with copyright risks is included, it is replaced with copyright-free music data. This audio replacement process utilizes speech synthesis services such as Amazon Polly. This avoids legal issues related to copyright when videos are distributed.

[0513] For example, if a user is live-streaming in the middle of town, it doesn't matter if passersby appear in the camera's view or if radio music is playing in the background. Because the server handles all processing in real time, the user can continue streaming without being aware of these issues.

[0514] An example of a prompt sentence to input into the generating AI model is, "Ideas for real-time video editing when a friend accidentally appears in a live stream with your pet."

[0515] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0516] Step 1:

[0517] The user records live video and audio using a communication terminal. Raw video and audio data from the camera and microphone are used as input. This data is streamed to the server in real time.

[0518] Step 2:

[0519] The server processes the received video data using image analysis technologies such as Amazon Recognition. The input is the raw video data sent in step 1. The server analyzes this data and extracts face information. The output is the state in which the face information has been identified.

[0520] Step 3:

[0521] The server compares the detected facial information with existing authorization information. Here, the facial information output in step 2 is used as input. If unauthorized facial information is included, the server selects alternative facial data from a pre-registered database. The output is a pair of the alternative facial data and the identified facial information.

[0522] Step 4:

[0523] The server uses compositing tools such as Unity to seamlessly composite the replacement face data onto the video data. The input for this step is the replacement face data and video data obtained in step 3. These are combined to generate video data that looks natural and doesn't feel unnatural. The output is the replaced video data.

[0524] Step 5:

[0525] The server analyzes the received audio data using AI technology. It analyzes this data using audio analysis technologies such as Amazon Polly. The input is the audio data transferred in step 1. The server detects the music portion that poses copyright risk from this data. The output is the audio data with the identified risk.

[0526] Step 6:

[0527] The server replaces audio data identified as having copyright risks with royalty-free music data. The audio data output in Step 5 and royalty-free music data obtained from the database are used as input. These are combined to generate a seamless audio track. The output is the replaced audio data.

[0528] Step 7:

[0529] The server combines the processed video and audio data and sends it back to the communication device. The input for this step is the final video and audio data after each processing step. The output is the processed live stream sent back to the user terminal, enabling real-time distribution.

[0530] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0531] This invention is a system that improves the user experience while avoiding issues related to portrait rights and copyrights by having a server, terminal, and user cooperate to analyze and process video and audio data within video content. In particular, by incorporating an emotion engine, it is possible to dynamically perform appropriate processing according to the context of the video.

[0532] In implementing the system, users first register the facial information they wish to allow via their device. This involves taking a photo of their face using the camera through a simple application operation, and uploading it to the server as part of the allow list. Furthermore, users can either upload videos to the server or configure live streaming settings.

[0533] The server analyzes the received video data and uses AI technology to recognize faces from the video frames. It then compares this face information with an allowed list, and if an unregistered face is found, it selects a random or situationally designated alternative face from the database and seamlessly integrates it into the video.

[0534] In addition, this system is equipped with an emotion engine that analyzes the user's facial expressions and gestures in the video to recognize their emotional state. Based on the recognized emotion, it dynamically changes the video's presentation—for example, music selection or face replacement—synchronizing the video's atmosphere with the emotion to provide a more natural and engaging viewing experience.

[0535] Furthermore, the server also analyzes the audio data in parallel to detect music that involves copyright. It replaces that portion with copyright-free music data stored in the database to generate a seamless audio track.

[0536] For example, if a user is smiling and live-tweeting an event, the emotion engine recognizes this cheerfulness and switches the background music to something bright and energetic. Conversely, if the user's emotions are calm, the music is also changed to a more appropriate, quiet melody. In this way, the system can harmonize the video and audio to match the user's psychological state through emotion recognition.

[0537] In this way, the system based on the present invention utilizes user permission data and emotion recognition to enable comprehensive management that ensures video content is emotionally engaging for viewers while complying with legal and ethical standards.

[0538] The following describes the processing flow.

[0539] Step 1:

[0540] Users use their devices to take photos of their own face and the faces of authorized individuals, registering them as a permission list. This data is uploaded from the device to the server.

[0541] Step 2:

[0542] The user either configures a live stream or uploads a recorded video file to the server. The server then begins preparing to analyze the received video data.

[0543] Step 3:

[0544] The server uses AI technology to analyze each frame of the video and detect human faces that appear in the footage. The detected face data is then compared with face information from a pre-submitted permission list.

[0545] Step 4:

[0546] The server identifies unauthorized facial information and selects alternative facial data from the database to replace it. This alternative face is processed using AI technology so that it is seamlessly integrated into the video.

[0547] Step 5:

[0548] The server activates an emotion engine and recognizes the user's emotional state by analyzing their facial expressions and gestures in the video.

[0549] Step 6:

[0550] The server selects and adjusts the music and surrogate face data used in the video based on the recognized emotional state. This synchronizes the mood of the video with the user's emotions.

[0551] Step 7:

[0552] The server simultaneously analyzes the audio data and detects copyrighted music. It then replaces that music with copyright-free music from its database to generate a natural-sounding audio track.

[0553] Step 8:

[0554] The server delivers the processed video to the user. The user can then stream or download this processed video in real time.

[0555] (Example 2)

[0556] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0557] Issues related to portrait rights and copyrights that arise when using video content prevent users from creating and sharing content with peace of mind. Furthermore, there is a need to effectively visualize emotional expressions within videos to improve the user experience. Additionally, traditional systems often provide static content that doesn't match the user's emotions.

[0558] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0559] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; and means for analyzing audio data and replacing identified copyright-risky music portions with copyright-free music data. This makes it possible to create video content that is engaging and synchronized with the user's emotions while avoiding issues related to portrait rights and copyright.

[0560] "Analyzing video data" means breaking down each frame of a video and processing it to identify specific elements.

[0561] "Identifying facial information" means extracting facial features detected from video data and comparing them with pre-registered data to check for a match.

[0562] "Selecting alternative face data" means that if an unauthorized face is present in the video, alternative face data will be extracted from the database.

[0563] "Analyzing audio data" refers to the process used to examine the audio track within a video and separate and identify specific music or audio components.

[0564] "Replacing with royalty-free music data" means changing the detected copyrighted music portion to music data that is free from licensing restrictions.

[0565] "Emotion recognition technology" is a technology that identifies a user's emotional state from their facial expressions and actions through image analysis.

[0566] "Adjusting the video's presentation" means dynamically changing elements within the video, such as music and visual effects, in response to the user's emotional state.

[0567] This invention provides a system that effectively avoids issues related to portrait rights and copyrights in video content and improves the user experience through the cooperation of a server, terminal, and user. In particular, by combining image analysis technology and emotion recognition technology and dynamically adjusting the content, it enables output that is more tailored to individual situations.

[0568] First, users register their facial information using their device. This is done using the photo-taking function installed on the device and through a facial recognition application. This image data is uploaded to the server and added to an allow list. Users who wish to process video data either upload the video data to the server via the internet or set up live streaming.

[0569] The server uses AI technology to analyze received video data and has the capability to recognize faces in the video frame by frame. During this process, face information within each frame is checked and compared against an allowed list. If face information is not registered, a substitute face is searched from the database and seamlessly integrated into the video data.

[0570] Emotion recognition technology analyzes users' facial expressions and actions in videos to recognize their emotional state. Based on the recognized emotional information, the server changes the video's presentation, specifically the background music and visual effects, according to the user. This synchronizes the video's atmosphere with the individual user's emotions, making it more engaging for viewers.

[0571] For example, consider a scenario where a user is filming their child's birthday party. The emotion engine recognizes bright smiles in the video and automatically changes the background music to a cheerful one to match the emotion. Similarly, if the scene is emotional, it adjusts the music to a calmer melody.

[0572] The generative AI model is given instructions using prompts. For example, specific instructions such as, "The user is filming their child's birthday party. The scene contains many smiles and celebratory gestures. Select music that suits this scene and appropriately process the unregistered facial information," are possible.

[0573] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0574] Step 1:

[0575] User facial information registration

[0576] The user takes a photo of their face using the device. This input is the captured image data. The device uploads the image to the server via the application and registers it in the allow list. Specifically, the user activates the device's camera and takes a photo of their face following the on-screen instructions. After taking the photo, the application automatically sends the image to the server. The output is the face data registered in the server's allow list.

[0577] Step 2:

[0578] Video upload or live stream settings

[0579] The user selects the video they want to analyze and either uploads it to the server or configures live streaming settings. The input is the video file or streaming information selected by the user. The device then sends this information to the server. Specifically, the user opens a file selection screen on their device, chooses a video, and presses the upload button. The output is the video data stored on the server.

[0580] Step 3:

[0581] Video data analysis and facial recognition

[0582] The server analyzes the received video data frame by frame using AI technology to recognize faces. The input is each frame of the video, and the output is the recognized face information and its location. Specifically, the AI ​​model scans each frame of the video and detects faces. The server then uses this information to prepare for the next step.

[0583] Step 4:

[0584] Facial information matching and surname face synthesis

[0585] The server compares the recognized facial information against an allow list. The input is the output of step 3. If an unregistered face is found as a result of the comparison, the server selects an appropriate substitute face from the database. Specifically, the server then performs a synthesis process and integrates the selected substitute face into the video. At this time, the substitute face is adjusted to match the facial expression and angle of the original video. The output is new video data with corrected facial information.

[0586] Step 5:

[0587] Emotion recognition and performance adjustment

[0588] The server uses emotion recognition technology to analyze the user's facial expressions and gestures in the video and identify their emotional state. The input consists of the original video data and face recognition data. The output is the recognized emotion information, which is used to change the video's music and effects. For example, if the server detects a smile from the user, it will switch to upbeat music that matches the smile. The specific operation involves facial emotion analysis by an AI model and automatic music selection from a music library.

[0589] Step 6:

[0590] Audio data analysis and music replacement

[0591] The server analyzes the audio track and detects music that may infringe on copyright. The input is the audio data of the video. If there is a copyright issue, the affected portion is replaced with copyright-free music. Specifically, the server scans the audio data, detects specific musical phrases, and replaces them. The output is a new video file containing the adjusted audio track.

[0592] (Application Example 2)

[0593] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0594] In the production and distribution of video content, there is a need to mitigate the risks of infringing on portrait rights and copyrights, while also creating emotionally engaging content for viewers. However, current technology lacks the systems to simultaneously satisfy these requirements. Specifically, there are risks such as individuals' faces appearing without permission, the use of copyrighted music, and the content's atmosphere not matching what viewers expect.

[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0596] In this invention, the server includes means for analyzing video information and identifying detected facial information by comparing it with pre-registered permission data; means for selecting alternative facial information to replace the identified facial information and naturally synthesizing it with the video information; means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information; and means for analyzing facial and movement information in the video using an emotion analysis function and dynamically changing the video and audio presentation according to the recognized emotional state. This makes it possible for content creators and distributors to produce and provide emotionally appealing video content to viewers while avoiding issues related to portrait rights and copyrights.

[0597] "Video information" refers to the visual data displayed on the screen, including video content and its individual frames.

[0598] "Facial information" refers to identifiable data about a person's face included in video information, including the position and characteristics of the face.

[0599] "Permission data" refers to information that has been registered in advance and confirmed to be permitted within the system, and specifically refers to facial information for which permission has been granted.

[0600] "Alternative facial information" refers to facial information used to replace facial information that does not match the permission data.

[0601] "Audio information" refers to digital data including voice and music, and specifically to sound information provided in conjunction with video information.

[0602] "Copyright-free music information" refers to music data that is not restricted by copyright and can be used freely.

[0603] "Emotional analysis function" refers to data and algorithms that analyze and infer an individual's emotional state from video and audio information.

[0604] "Dynamic modification" refers to the real-time editing and reconfiguration of video and audio data in response to specific conditions.

[0605] The system for realizing an application of this invention operates through the cooperation of three parties: a server, a terminal, and a user. The user uses the terminal to record video and uploads it to the server. The terminal uses a camera application with facial recognition capabilities to capture facial information authorized by the user and registers it with the server as authorized data in advance.

[0606] The server uses deep learning technology to identify faces in order to analyze the received video information. Specifically, it analyzes the video information frame by frame, compares the detected faces with permission data, selects alternative faces from the database for unregistered faces, and synthesizes them naturally using libraries such as OpenCV. This process makes it possible to avoid issues related to portrait rights.

[0607] Furthermore, the server analyzes the audio information to identify copyright-infringing musical portions. The FFmpeg library is used to replace these detected portions with copyright-free music. During this process, AWS cloud services with user emotion analysis capabilities are utilized to analyze the user's facial expressions and actions within the video. This emotion analysis technology enables the video and audio to be dynamically modified in real time according to the recognized emotional state.

[0608] For example, if a user uploads a video they filmed while traveling to the server, the server will recognize the user's smile and automatically insert upbeat background music that matches the video. In this way, it is possible to provide emotionally appealing content to viewers.

[0609] An example of a prompt might be, "Please create a video with many smiling user images and add cheerful background music." This prompt can be used as input for a generative AI model.

[0610] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0611] Step 1:

[0612] The user records a video using their device. During this process, the user uses a camera application to record the video and generate video data. The recorded video is immediately saved to the device, and preparations for facial recognition to extract facial information are also initiated within the same application.

[0613] Step 2:

[0614] The user registers the facial information they wish to allow using their device. Specifically, they take a picture of their face using an application with facial recognition capabilities and store it as permission data. The input is facial image data, and the output is a list of permission data sent to the server.

[0615] Step 3:

[0616] The user uploads the captured video information from their device to the server. The device retrieves the video data and sends it to the server via the network. The input is the video data, and the output is the video information stored on the server.

[0617] Step 4:

[0618] The server analyzes the received video information and performs face recognition. It uses a deep learning algorithm to detect face information in each frame. The input is the video data on the server, and the output is a list of detected face information.

[0619] Step 5:

[0620] The server compares the detected face information with the permission data list and generates alternative face information for unregistered faces. If no match is found, information is selected from the alternative face database and synthesized into the video data using OpenCV or similar tools. The output is new video data with the alternative face applied.

[0621] Step 6:

[0622] The server analyzes the audio information and identifies the music portions that pose copyright risks. After identifying the risky portions through audio analysis, it performs replacement processing using the FFmpeg library. The input is audio data on the server, and the output is audio data with royalty-free music set as the input.

[0623] Step 7:

[0624] The server uses emotion analysis capabilities to analyze facial expressions and motion information within the video. It uses services such as AWS Rekognition to estimate the user's emotions. The input is video data on the server, and the output is the estimated emotional state.

[0625] Step 8:

[0626] The server dynamically changes the video and music effects according to the recognized emotional state. For example, if a smile is detected, the background music will be changed to a bright and energetic one. The input is the estimated emotional state and royalty-free music data, and the output is the final video data with the effects changed.

[0627] Through these steps, content creators can avoid legal risks while providing emotionally engaging video content for viewers.

[0628] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0629] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0630] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0631] [Fourth Embodiment]

[0632] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0633] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0634] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0635] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0636] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0637] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0638] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0639] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0640] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0641] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0642] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0643] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0644] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] The system for implementing this invention involves a server and a terminal working in cooperation. First, the user takes photos of their own face and the faces of authorized acquaintances using the terminal and sends them to the server as an authorized list. This face information is later stored as reference information used during video processing.

[0646] Next, the user uploads a live stream or recorded video to the server. Upon receiving the video data, the server uses AI technology to automatically extract facial information from each frame of the video. This allows the faces of people appearing in the video to be identified in real time and matched against permission information.

[0647] If, after comparing facial information, a face not present in the allowlist is detected, the server accesses the database and selects appropriate alternative facial data randomly or based on user specifications. The selected face is then synthesized using AI to blend naturally into the video, resulting in no visual inconsistencies.

[0648] The server also analyzes the audio data obtained from the video to identify music that may cause copyright issues. If detected, the server seamlessly replaces the relevant portion with royalty-free music data to generate a seamless audio track.

[0649] For example, when a user is live-streaming in the city, if a passerby unexpectedly appears in the frame, their face will be replaced in real time with another face, allowing the stream to continue. Similarly, if background music playing in a cafe or similar location is copyrighted, it will be automatically replaced with different, licensed music. This system performs all processing in the background, allowing users to stream videos with peace of mind without having to go through complicated procedures.

[0650] The following describes the processing flow.

[0651] Step 1:

[0652] Users use their devices to take photos of themselves and authorized individuals, registering them as facial recognition data. This data is uploaded from the device to the server as a permission list.

[0653] Step 2:

[0654] Users upload live streams or previously recorded videos to the server. The server stores the received video files for analysis.

[0655] Step 3:

[0656] The server uses AI technology to analyze each frame of the video and detect faces. The detected face data is then compared and matched with the face information in the previously received permission list.

[0657] Step 4:

[0658] The server identifies faces not on the allowed list based on the matching results and selects alternative face data from the database to replace them. The selected alternative face is then seamlessly integrated into the corresponding portion of the video using AI.

[0659] Step 5:

[0660] The server extracts audio data from the video and analyzes the waveform patterns of the music to detect music that poses a copyright risk. If detected, the server selects appropriate copyright-free music from its database.

[0661] Step 6:

[0662] The server replaces the original music portion with detected, royalty-free music data, generating a natural-sounding audio track while maintaining the overall harmony of the audio track.

[0663] Step 7:

[0664] The server provides the processed video data to the user. The user can then securely distribute or download this video.

[0665] (Example 1)

[0666] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0667] Protecting privacy and avoiding copyright issues are crucial challenges in modern video distribution. In particular, the unintended appearance of faces or the use of background music can lead to the leakage of personal information and legal troubles. This invention aims to solve these problems.

[0668] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0669] In this invention, the server includes means for analyzing video information and identifying detected face identification information by comparing it with pre-registered permission information; means for selecting alternative face information to replace the identified face identification information and naturally synthesizing it with the video information; and means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information. This makes it possible to protect privacy and avoid copyright issues.

[0670] "Visual information" refers to visual data recorded in the form of images or videos, and this includes image data for each frame.

[0671] "Facial identification information" refers to feature data related to a person's face extracted from video information, and includes feature points such as the position of the eyes, nose, and mouth.

[0672] "Permission information" refers to a list of pre-registered facial identification information, which is used to match facial identification information within video footage.

[0673] "Alternative facial information" refers to artificial or other facial data selected to replace existing facial identification information within video data.

[0674] "Synthesizing" refers to the process of integrating multiple data sets and reconstructing them as a single, natural-looking visual data set.

[0675] "Audio information" refers to audio data accompanying video information, which includes conversations, music, and other sounds.

[0676] "Copyright-free music information" refers to music data that is not subject to copyright restrictions and can be freely used.

[0677] This invention utilizes a system in which a server and a terminal work in cooperation. First, the user acquires facial identification information using the terminal's camera. This information is captured using a dedicated application or camera function on the terminal and transmitted to an information processing device in order to be included in a destination list. The server stores this received facial identification information as permission information in its data storage.

[0678] The user then uploads the live stream or recorded video information to the server. The server processes this video information, using software such as TensorFlow and OpenCV. These tools are used to analyze the video frame by frame and automatically extract face identification information. This allows for the identification of people's faces based on permission information.

[0679] If the identified face is not present in the permission information, the server consults the database and selects alternative face information. The selected face information is then seamlessly integrated into the video information using DeepFake technology or StyleGAN. For example, if a passerby unexpectedly appears in a live stream of a city street, the passerby's face is replaced with another face in real time.

[0680] Furthermore, the server also analyzes audio information to identify music that poses copyright risks. Using AI-powered speech recognition, the music is replaced with royalty-free music. For example, background music playing when filming in a cafe or similar location is automatically replaced with licensed music.

[0681] Here's an example of a specific prompt when using a generative AI model: "I want to develop a system that replaces specific faces in real time with other faces in live-streamed video from a city street. Please tell me what essential technologies are needed. Also, please advise on how to automatically replace specific background music within an audio track."

[0682] This system allows users to protect their privacy, avoid copyright risks, and achieve safe and smooth video streaming.

[0683] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0684] Step 1:

[0685] The user uses the device's camera to acquire facial recognition information. The input is a facial image of the user or an authorized acquaintance. The device uses a facial recognition library such as OpenCV to extract facial feature points. The output is digital data ready to be sent to the server as facial recognition information. Specifically, it provides numerical data representing the positions of the eyes, nose, and mouth.

[0686] Step 2:

[0687] The terminal sends the extracted facial identification information to the server. The input is facial identification information data, and the output is a list of permission information stored in the server's database. The server receives this data using a secure communication protocol (e.g., HTTPS) and stores it in storage to use it as reference information. Specifically, the facial data is encrypted before transmission, and the server verifies data integrity before storing it.

[0688] Step 3:

[0689] The user uploads live stream or recorded video information to the server. The input is a video file captured on the user's device. The server divides this video data frame by frame and uses TensorFlow or OpenCV to extract face identification information from each frame. The output is face identification information for each frame, which is used in the subsequent face matching process. Specifically, it is processed sequentially according to the frame rate, and data is generated in which face regions are identified and extracted.

[0690] Step 4:

[0691] The server compares the extracted face identification information with the permission information to perform identification. The input is the face identification information for each frame, and the output is the result of determining whether it matches or does not match the permission information. A database query is used for the comparison, and as a result, face identification information that does not exist in the permission information is identified. Specifically, the comparison algorithm is used to improve the accuracy of face recognition.

[0692] Step 5:

[0693] If the server detects a face that does not exist in the permission information, it selects alternative face information and composites it into the video. The input is the face identification information that was determined to be a mismatch, and the output is the video with the alternative face composited. DeepFake technology and StyleGAN are used for the synthesis, and the selected alternative face information is integrated into the video in real time. Specifically, the process involves generating different face data and applying it to the video frames in a natural way.

[0694] Step 6:

[0695] The server analyzes the audio information of the video to identify copyright-infringing musical portions. The input is the audio data contained in the video, and the output is audio data with copyright-free music replaced. AI-based music fingerprinting technology is used for the analysis, and the detected musical portions are appropriately replaced. Specifically, the audio track is split, and the problematic portions undergo an overwrite process, including volume adjustment.

[0696] This allows users to distribute videos free from privacy and copyright concerns.

[0697] (Application Example 1)

[0698] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0699] In modern content distribution services, protecting privacy and managing copyrights are critical issues. In particular, the frequent occurrence of unintended individuals appearing on live streams or copyrighted music playing in the background poses legal risks for both streamers and platforms. Addressing these challenges and ensuring compliance while maintaining freedom of distribution is essential.

[0700] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0701] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; means for analyzing audio data and replacing identified copyright-risk music portions with copyright-free music data; and means for processing video and audio data in real time and transmitting it to a communication device. This enables distributors to safely deliver content in real time without concerns about privacy or copyright issues.

[0702] "Video data" refers to visual information acquired from a camera or other recording device, and is composed of multiple frames.

[0703] "Analysis" is the process of breaking down data into smaller parts and revealing its constituent elements and functions.

[0704] "Facial information" refers to data that includes facial feature points and patterns used to identify individuals.

[0705] "Permission information" refers to data and information that has been registered in advance to confirm that specific conditions or criteria are met.

[0706] "Substitution" is the operation of replacing one element with another.

[0707] "Alternative face data" refers to data about another face that is used in place of the original face information, and is designed to be synthesized naturally.

[0708] "Natural synthesis" refers to techniques that ensure that no visual inconsistencies occur when modifications or substitutions are made.

[0709] "Audio data" refers to auditory information obtained from sound acquisition devices such as microphones.

[0710] "Music portions with copyright risks" refer to segments of music that are protected by copyright and whose use without permission may cause legal problems.

[0711] "Copyright-free music data" refers to music data for which the right to freely use under specific usage conditions has been granted.

[0712] "Real-time" refers to a process where data acquisition, processing, and output are synchronized with the passage of time in the real world.

[0713] A "communication device" is a device or equipment that sends and receives data and has the ability to transmit information over a network.

[0714] The system that realizes this invention operates with a server and a terminal working together. It starts when a user uses a smartphone or other communication terminal to record live video and sends the data to the server. In this process, the terminal acquires video and audio data in real time from the camera and microphone.

[0715] The server analyzes video data using AI technologies such as Amazon Recognition, detects facial information, and compares it with pre-registered permission information. If unauthorized facial information is included, the server selects alternative facial data from the database and uses software such as Unity to seamlessly composite it into the video data. This ensures that even if a specific individual is captured in the video, their face is replaced with another face without any visual incongruity.

[0716] For audio data, AI technology is used to analyze identified musical portions, and if music with copyright risks is included, it is replaced with copyright-free music data. This audio replacement process utilizes speech synthesis services such as Amazon Polly. This avoids legal issues related to copyright when videos are distributed.

[0717] For example, if a user is live-streaming in the middle of town, it doesn't matter if passersby appear in the camera's view or if radio music is playing in the background. Because the server handles all processing in real time, the user can continue streaming without being aware of these issues.

[0718] An example of a prompt sentence to input into the generating AI model is, "Ideas for real-time video editing when a friend accidentally appears in a live stream with your pet."

[0719] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0720] Step 1:

[0721] The user records live video and audio using a communication terminal. Raw video and audio data from the camera and microphone are used as input. This data is streamed to the server in real time.

[0722] Step 2:

[0723] The server processes the received video data using image analysis technologies such as Amazon Recognition. The input is the raw video data sent in step 1. The server analyzes this data and extracts face information. The output is the state in which the face information has been identified.

[0724] Step 3:

[0725] The server compares the detected facial information with existing authorization information. Here, the facial information output in step 2 is used as input. If unauthorized facial information is included, the server selects alternative facial data from a pre-registered database. The output is a pair of the alternative facial data and the identified facial information.

[0726] Step 4:

[0727] The server uses compositing tools such as Unity to seamlessly composite the replacement face data onto the video data. The input for this step is the replacement face data and video data obtained in step 3. These are combined to generate video data that looks natural and doesn't feel unnatural. The output is the replaced video data.

[0728] Step 5:

[0729] The server analyzes the received audio data using AI technology. It analyzes this data using audio analysis technologies such as Amazon Polly. The input is the audio data transferred in step 1. The server detects the music portion that poses copyright risk from this data. The output is the audio data with the identified risk.

[0730] Step 6:

[0731] The server replaces audio data identified as having copyright risks with royalty-free music data. The audio data output in Step 5 and royalty-free music data obtained from the database are used as input. These are combined to generate a seamless audio track. The output is the replaced audio data.

[0732] Step 7:

[0733] The server combines the processed video and audio data and sends it back to the communication device. The input for this step is the final video and audio data after each processing step. The output is the processed live stream sent back to the user terminal, enabling real-time distribution.

[0734] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0735] This invention is a system that improves the user experience while avoiding issues related to portrait rights and copyrights by having a server, terminal, and user cooperate to analyze and process video and audio data within video content. In particular, by incorporating an emotion engine, it is possible to dynamically perform appropriate processing according to the context of the video.

[0736] In implementing the system, users first register the facial information they wish to allow via their device. This involves taking a photo of their face using the camera through a simple application operation, and uploading it to the server as part of the allow list. Furthermore, users can either upload videos to the server or configure live streaming settings.

[0737] The server analyzes the received video data and uses AI technology to recognize faces from the video frames. It then compares this face information with an allowed list, and if an unregistered face is found, it selects a random or situationally designated alternative face from the database and seamlessly integrates it into the video.

[0738] In addition, this system is equipped with an emotion engine that analyzes the user's facial expressions and gestures in the video to recognize their emotional state. Based on the recognized emotion, it dynamically changes the video's presentation—for example, music selection or face replacement—synchronizing the video's atmosphere with the emotion to provide a more natural and engaging viewing experience.

[0739] Furthermore, the server also analyzes the audio data in parallel to detect music that involves copyright. It replaces that portion with copyright-free music data stored in the database to generate a seamless audio track.

[0740] For example, if a user is smiling and live-tweeting an event, the emotion engine recognizes this cheerfulness and switches the background music to something bright and energetic. Conversely, if the user's emotions are calm, the music is also changed to a more appropriate, quiet melody. In this way, the system can harmonize the video and audio to match the user's psychological state through emotion recognition.

[0741] In this way, the system based on the present invention utilizes user permission data and emotion recognition to enable comprehensive management that ensures video content is emotionally engaging for viewers while complying with legal and ethical standards.

[0742] The following describes the processing flow.

[0743] Step 1:

[0744] Users use their devices to take photos of their own face and the faces of authorized individuals, registering them as a permission list. This data is uploaded from the device to the server.

[0745] Step 2:

[0746] The user either configures a live stream or uploads a recorded video file to the server. The server then begins preparing to analyze the received video data.

[0747] Step 3:

[0748] The server uses AI technology to analyze each frame of the video and detect human faces that appear in the footage. The detected face data is then compared with face information from a pre-submitted permission list.

[0749] Step 4:

[0750] The server identifies unauthorized facial information and selects alternative facial data from the database to replace it. This alternative face is processed using AI technology so that it is seamlessly integrated into the video.

[0751] Step 5:

[0752] The server activates an emotion engine and recognizes the user's emotional state by analyzing their facial expressions and gestures in the video.

[0753] Step 6:

[0754] The server selects and adjusts the music and surrogate face data used in the video based on the recognized emotional state. This synchronizes the mood of the video with the user's emotions.

[0755] Step 7:

[0756] The server simultaneously analyzes the audio data and detects copyrighted music. It then replaces that music with copyright-free music from its database to generate a natural-sounding audio track.

[0757] Step 8:

[0758] The server delivers the processed video to the user. The user can then stream or download this processed video in real time.

[0759] (Example 2)

[0760] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0761] Issues related to portrait rights and copyrights that arise when using video content prevent users from creating and sharing content with peace of mind. Furthermore, there is a need to effectively visualize emotional expressions within videos to improve the user experience. Additionally, traditional systems often provide static content that doesn't match the user's emotions.

[0762] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0763] In this invention, the server includes means for analyzing video data and identifying detected facial information by comparing it with pre-registered permission information; means for selecting alternative facial data to replace the identified facial information and naturally synthesizing it with the video data; and means for analyzing audio data and replacing identified copyright-risky music portions with copyright-free music data. This makes it possible to create video content that is engaging and synchronized with the user's emotions while avoiding issues related to portrait rights and copyright.

[0764] "Analyzing video data" means breaking down each frame of a video and processing it to identify specific elements.

[0765] "Identifying facial information" means extracting facial features detected from video data and comparing them with pre-registered data to check for a match.

[0766] "Selecting alternative face data" means that if an unauthorized face is present in the video, alternative face data will be extracted from the database.

[0767] "Analyzing audio data" refers to the process used to examine the audio track within a video and separate and identify specific music or audio components.

[0768] "Replacing with royalty-free music data" means changing the detected copyrighted music portion to music data that is free from licensing restrictions.

[0769] "Emotion recognition technology" is a technology that identifies a user's emotional state from their facial expressions and actions through image analysis.

[0770] "Adjusting the video's presentation" means dynamically changing elements within the video, such as music and visual effects, in response to the user's emotional state.

[0771] This invention provides a system that effectively avoids issues related to portrait rights and copyrights in video content and improves the user experience through the cooperation of a server, terminal, and user. In particular, by combining image analysis technology and emotion recognition technology and dynamically adjusting the content, it enables output that is more tailored to individual situations.

[0772] First, users register their facial information using their device. This is done using the photo-taking function installed on the device and through a facial recognition application. This image data is uploaded to the server and added to an allow list. Users who wish to process video data either upload the video data to the server via the internet or set up live streaming.

[0773] The server uses AI technology to analyze received video data and has the capability to recognize faces in the video frame by frame. During this process, face information within each frame is checked and compared against an allowed list. If face information is not registered, a substitute face is searched from the database and seamlessly integrated into the video data.

[0774] Emotion recognition technology analyzes users' facial expressions and actions in videos to recognize their emotional state. Based on the recognized emotional information, the server changes the video's presentation, specifically the background music and visual effects, according to the user. This synchronizes the video's atmosphere with the individual user's emotions, making it more engaging for viewers.

[0775] For example, consider a scenario where a user is filming their child's birthday party. The emotion engine recognizes bright smiles in the video and automatically changes the background music to a cheerful one to match the emotion. Similarly, if the scene is emotional, it adjusts the music to a calmer melody.

[0776] The generative AI model is given instructions using prompts. For example, specific instructions such as, "The user is filming their child's birthday party. The scene contains many smiles and celebratory gestures. Select music that suits this scene and appropriately process the unregistered facial information," are possible.

[0777] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0778] Step 1:

[0779] User facial information registration

[0780] The user takes a photo of their face using the device. This input is the captured image data. The device uploads the image to the server via the application and registers it in the allow list. Specifically, the user activates the device's camera and takes a photo of their face following the on-screen instructions. After taking the photo, the application automatically sends the image to the server. The output is the face data registered in the server's allow list.

[0781] Step 2:

[0782] Video upload or live stream settings

[0783] The user selects the video they want to analyze and either uploads it to the server or configures live streaming settings. The input is the video file or streaming information selected by the user. The device then sends this information to the server. Specifically, the user opens a file selection screen on their device, chooses a video, and presses the upload button. The output is the video data stored on the server.

[0784] Step 3:

[0785] Video data analysis and facial recognition

[0786] The server analyzes the received video data frame by frame using AI technology to recognize faces. The input is each frame of the video, and the output is the recognized face information and its location. Specifically, the AI ​​model scans each frame of the video and detects faces. The server then uses this information to prepare for the next step.

[0787] Step 4:

[0788] Facial information matching and surname face synthesis

[0789] The server compares the recognized facial information against an allow list. The input is the output of step 3. If an unregistered face is found as a result of the comparison, the server selects an appropriate substitute face from the database. Specifically, the server then performs a synthesis process and integrates the selected substitute face into the video. At this time, the substitute face is adjusted to match the facial expression and angle of the original video. The output is new video data with corrected facial information.

[0790] Step 5:

[0791] Emotion recognition and performance adjustment

[0792] The server uses emotion recognition technology to analyze the user's facial expressions and gestures in the video and identify their emotional state. The input consists of the original video data and face recognition data. The output is the recognized emotion information, which is used to change the video's music and effects. For example, if the server detects a smile from the user, it will switch to upbeat music that matches the smile. The specific operation involves facial emotion analysis by an AI model and automatic music selection from a music library.

[0793] Step 6:

[0794] Audio data analysis and music replacement

[0795] The server analyzes the audio track and detects music that may infringe on copyright. The input is the audio data of the video. If there is a copyright issue, the affected portion is replaced with copyright-free music. Specifically, the server scans the audio data, detects specific musical phrases, and replaces them. The output is a new video file containing the adjusted audio track.

[0796] (Application Example 2)

[0797] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0798] In the production and distribution of video content, there is a need to mitigate the risks of infringing on portrait rights and copyrights, while also creating emotionally engaging content for viewers. However, current technology lacks the systems to simultaneously satisfy these requirements. Specifically, there are risks such as individuals' faces appearing without permission, the use of copyrighted music, and the content's atmosphere not matching what viewers expect.

[0799] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0800] In this invention, the server includes means for analyzing video information and identifying detected facial information by comparing it with pre-registered permission data; means for selecting alternative facial information to replace the identified facial information and naturally synthesizing it with the video information; means for analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information; and means for analyzing facial and movement information in the video using an emotion analysis function and dynamically changing the video and audio presentation according to the recognized emotional state. This makes it possible for content creators and distributors to produce and provide emotionally appealing video content to viewers while avoiding issues related to portrait rights and copyrights.

[0801] "Video information" refers to the visual data displayed on the screen, including video content and its individual frames.

[0802] "Facial information" refers to identifiable data about a person's face included in video information, including the position and characteristics of the face.

[0803] "Permission data" refers to information that has been registered in advance and confirmed to be permitted within the system, and specifically refers to facial information for which permission has been granted.

[0804] "Alternative facial information" refers to facial information used to replace facial information that does not match the permission data.

[0805] "Audio information" refers to digital data including voice and music, and specifically to sound information provided in conjunction with video information.

[0806] "Copyright-free music information" refers to music data that is not restricted by copyright and can be used freely.

[0807] "Emotional analysis function" refers to data and algorithms that analyze and infer an individual's emotional state from video and audio information.

[0808] "Dynamic modification" refers to the real-time editing and reconfiguration of video and audio data in response to specific conditions.

[0809] The system for realizing an application of this invention operates through the cooperation of three parties: a server, a terminal, and a user. The user uses the terminal to record video and uploads it to the server. The terminal uses a camera application with facial recognition capabilities to capture facial information authorized by the user and registers it with the server as authorized data in advance.

[0810] The server uses deep learning technology to identify faces in order to analyze the received video information. Specifically, it analyzes the video information frame by frame, compares the detected faces with permission data, selects alternative faces from the database for unregistered faces, and synthesizes them naturally using libraries such as OpenCV. This process makes it possible to avoid issues related to portrait rights.

[0811] Furthermore, the server analyzes the audio information to identify copyright-infringing musical portions. The FFmpeg library is used to replace these detected portions with copyright-free music. During this process, AWS cloud services with user emotion analysis capabilities are utilized to analyze the user's facial expressions and actions within the video. This emotion analysis technology enables the video and audio to be dynamically modified in real time according to the recognized emotional state.

[0812] For example, if a user uploads a video they filmed while traveling to the server, the server will recognize the user's smile and automatically insert upbeat background music that matches the video. In this way, it is possible to provide emotionally appealing content to viewers.

[0813] An example of a prompt might be, "Please create a video with many smiling user images and add cheerful background music." This prompt can be used as input for a generative AI model.

[0814] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0815] Step 1:

[0816] The user records a video using their device. During this process, the user uses a camera application to record the video and generate video data. The recorded video is immediately saved to the device, and preparations for facial recognition to extract facial information are also initiated within the same application.

[0817] Step 2:

[0818] The user registers the facial information they wish to allow using their device. Specifically, they take a picture of their face using an application with facial recognition capabilities and store it as permission data. The input is facial image data, and the output is a list of permission data sent to the server.

[0819] Step 3:

[0820] The user uploads the captured video information from their device to the server. The device retrieves the video data and sends it to the server via the network. The input is the video data, and the output is the video information stored on the server.

[0821] Step 4:

[0822] The server analyzes the received video information and performs face recognition. It uses a deep learning algorithm to detect face information in each frame. The input is the video data on the server, and the output is a list of detected face information.

[0823] Step 5:

[0824] The server compares the detected face information with the permission data list and generates alternative face information for unregistered faces. If no match is found, information is selected from the alternative face database and synthesized into the video data using OpenCV or similar tools. The output is new video data with the alternative face applied.

[0825] Step 6:

[0826] The server analyzes the audio information and identifies the music portions that pose copyright risks. After identifying the risky portions through audio analysis, it performs replacement processing using the FFmpeg library. The input is audio data on the server, and the output is audio data with royalty-free music set as the input.

[0827] Step 7:

[0828] The server uses emotion analysis capabilities to analyze facial expressions and motion information within the video. It uses services such as AWS Rekognition to estimate the user's emotions. The input is video data on the server, and the output is the estimated emotional state.

[0829] Step 8:

[0830] The server dynamically changes the video and music effects according to the recognized emotional state. For example, if a smile is detected, the background music will be changed to a bright and energetic one. The input is the estimated emotional state and royalty-free music data, and the output is the final video data with the effects changed.

[0831] Through these steps, content creators can avoid legal risks while providing emotionally engaging video content for viewers.

[0832] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0833] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0834] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0835] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0836] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0837] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0838] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0839] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0840] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0841] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0842] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0843] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0844] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0845] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0846] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0847] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0848] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0849] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0850] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0851] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0852] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0853] The following is further disclosed regarding the embodiments described above.

[0854] (Claim 1)

[0855] A means for analyzing video data and identifying detected facial information by matching it with pre-registered permission information,

[0856] A means for selecting alternative facial data to replace identified facial information and naturally compositing it into video data,

[0857] A method for analyzing audio data and replacing identified copyright-risk musical portions with copyright-free music data,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, wherein the user registers their permitted facial information in advance and transmits it to the server.

[0861] (Claim 3)

[0862] The system according to claim 1, which provides processed video data to the user in the form of distribution or download.

[0863] "Example 1"

[0864] (Claim 1)

[0865] A means for analyzing video information and identifying the detected face by comparing it with pre-registered permission information,

[0866] A means for selecting alternative facial information to replace identified facial identification information and naturally synthesizing it with video information,

[0867] A means of analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information,

[0868] A system that includes this.

[0869] (Claim 2)

[0870] The system according to claim 1, which registers user-authorized facial identification information in advance and transmits it to an information processing device.

[0871] (Claim 3)

[0872] The system according to claim 1, which provides processed video information to a user in the form of distribution or acquisition.

[0873] "Application Example 1"

[0874] (Claim 1)

[0875] A means for analyzing video data and identifying detected facial information by matching it with pre-registered permission information,

[0876] A means for selecting alternative facial data to replace identified facial information and naturally compositing it into video data,

[0877] A method for analyzing audio data and replacing identified copyright-risk musical portions with copyright-free music data,

[0878] A means for processing video and audio data in real time and transmitting it to a communication device,

[0879] A system that includes this.

[0880] (Claim 2)

[0881] The system according to claim 1, wherein the user registers their permitted facial information in advance and transmits it to the server.

[0882] (Claim 3)

[0883] The system according to claim 1, which distributes processed video data to users or provides it to users in the form of an information recording medium.

[0884] "Example 2 of combining an emotion engine"

[0885] (Claim 1)

[0886] A means for analyzing video data and identifying detected facial information by matching it with pre-registered permission information,

[0887] A means for selecting alternative facial data to replace identified facial information and naturally compositing it into video data,

[0888] A method for analyzing audio data and replacing identified copyright-risk musical portions with copyright-free music data,

[0889] A means of analyzing the user's emotional state using emotion recognition technology and dynamically adjusting the presentation of video data based on that emotion,

[0890] A system that includes this.

[0891] (Claim 2)

[0892] The system according to claim 1, wherein the user registers their permitted facial information in advance and transmits it to the server.

[0893] (Claim 3)

[0894] The system according to claim 1, which provides processed video data to the user in the form of distribution or download.

[0895] "Application example 2 when combining with an emotional engine"

[0896] (Claim 1)

[0897] A means for analyzing video information and identifying detected facial information by comparing it with pre-registered permission data,

[0898] A means for selecting alternative facial information to replace identified facial information and naturally synthesizing it with video information,

[0899] A means of analyzing audio information and replacing identified copyright-risk music portions with copyright-free music information,

[0900] A means for analyzing facial expressions and movement information in a video using an emotion analysis function, and dynamically changing the video and audio effects according to the recognized emotional state,

[0901] A system that includes this.

[0902] (Claim 2)

[0903] The system according to claim 1, which pre-registers facial information authorized by the user and transmits it to an information processing device.

[0904] (Claim 3)

[0905] The system according to claim 1, which provides processed video information to the user in the form of distribution or download. [Explanation of symbols]

[0906] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for analyzing video data and identifying detected facial information by matching it with pre-registered permission information, A means for selecting alternative facial data to replace identified facial information and naturally compositing it into video data, A method for analyzing audio data and replacing identified copyright-risk musical portions with copyright-free music data, A means for processing video and audio data in real time and transmitting it to a communication device, A system that includes this.

2. The system according to claim 1, wherein the user registers their permitted facial information in advance and transmits it to the server.

3. The system according to claim 1, which distributes processed video data to users or provides it to users in the form of an information recording medium.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A