System
The system addresses the lack of personalization and immersion in movie and video content by using real-time face and voice capture with generative AI to replace characters, ensuring a secure and personalized viewing experience.
Patent Information
- Application Number
- JP2024119132
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing movie and video content viewing experiences lack personalization and immersion, particularly for children and fans of their favorite idols, and there is a need for secure controls to prevent inappropriate content generation.
A system that captures a user's face and voice in real time, uses generative AI to replace characters' faces and voices with the user's, and securely transmits and regenerates the content for personalized viewing, using a secure communication protocol and compression to ensure safety.
Enables a more personal and immersive viewing experience while preventing the creation of inappropriate content by securely customizing movie and video content with the user's face and voice.
Smart Images

Figure 2026018071000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] There are challenges in the viewing experience of movies and video content, making it difficult for viewers to enjoy a more personal and immersive experience. This challenge is particularly pronounced for children and fans of their favorite idols, who expect a deep emotional connection and the ability to identify with the content. There is also a need for controls and security to prevent the creation of content that is inappropriate for others. [Means for solving the problem]
[0005] This invention is a system that captures a user's face and voice in real time, inputs this data into a generation AI, and replaces the face and voice of characters in content with the user's face and voice. The captured data is sent to a server using a secure communication protocol, and the content regenerated based on the input data is compressed, encoded, and sent to the user's device. This allows users to be more immersed in the content and enjoy a more personal experience, while also providing security to prevent the generation of inappropriate content.
[0006] A "user" is a person using the system to personally customize their viewing experience of movie or video content.
[0007] "Face" is video data of the user's face, and is used to replace the face of a character in the content.
[0008] "Voice" is data of the user's voice, and is used to replace the voice of a character in the content.
[0009] "Capture" refers to the process of acquiring video and audio in real time and saving them as digital data.
[0010] "Generative AI" refers to artificial intelligence that generates new images and sounds based on the input of a user's facial and voice data.
[0011] "Input" refers to the process of entering digital data into a system or algorithm.
[0012] "Content" refers to viewable multimedia data such as movies and videos.
[0013] "Characters" refers to characters who appear in movies or video content and have roles.
[0014] "Replace" refers to the process of changing the original video and audio data to the user's data.
[0015] "Regeneration" refers to the process of rebuilding content based on new data.
[0016] A "secure communication protocol" refers to a communication method that ensures safety when sending and receiving data.
[0017] "Compression" refers to the process of reducing the size of digital data.
[0018] "Encoding" refers to the process of converting digital data into a particular format.
[0019] A "terminal" is a device that is directly operated by a user and is used for capturing, playing back content, and communicating. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] The present invention is a system for personalizing the viewing experience of movies and video content. Users can use their own face and voice to replace characters in the content with themselves. This system is characterized by exchanging data using a secure communication protocol and playing the generated customized content on the user's terminal. Specific embodiments of the system are described below.
[0042] User operations
[0043] 1. Launching the application
[0044] The user launches an application on the device by tapping the app icon on the device's home screen.
[0045] 2. Face and voice capture
[0046] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0047] When a user faces the camera and speaks into the microphone for a few seconds, the device captures and records video and audio in real time.
[0048] 3. Content Selection
[0049] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[0050] Server Processing
[0051] 1. Data Receipt and Analysis
[0052] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[0053] 2. Acquiring character data
[0054] The server retrieves character data (face, voice, movements, lines) from the database of the selected content.
[0055] 3. Use of generative AI
[0056] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the inputted face and voice data of the user.
[0057] 4. Generating new lines
[0058] The server changes the lines that other characters use to call out to the user to new lines that include the user's name, so that other characters will call out the user's name.
[0059] 5. Building regenerative content
[0060] The server reconstructs the original video content based on the generated facial, voice, and dialogue data, and then uses editing software to integrate the new characters and voices into the original data.
[0061] 6. Data Compression and Transfer
[0062] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[0063] Terminal handling
[0064] 1. Receiving new content
[0065] The terminal receives the regenerated content sent from the server and stores it in local storage.
[0066] 2. Playing content
[0067] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0068] Specific examples
[0069] For example, suppose Person A captures his or her own face and voice and selects an animated movie. The server inputs Person A's face and voice into the generation AI, which then replaces the face and voice of the main character in the animated movie with Person A's. It also changes the lines of other characters to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie starring himself or herself as the main character.
[0070] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent the generation of inappropriate content.
[0071] The processing flow will be explained below.
[0072] Step 1:
[0073] The user launches an application on the device by tapping the corresponding app icon on the device's home screen.
[0074] Step 2:
[0075] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0076] Step 3:
[0077] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[0078] Step 4:
[0079] The user selects the movie or video content they want to watch through the device's interface, and the device selects the title from its content library.
[0080] Step 5:
[0081] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[0082] Step 6:
[0083] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0084] Step 7:
[0085] The server uses generative AI to replace the face and voice of the characters in the content with those of the user based on the inputted face and voice data of the user, and generates a digital model that reflects the user's characteristics.
[0086] Step 8:
[0087] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0088] Step 9:
[0089] The server recreates the original video content based on the generated facial, voice, and dialogue data, and uses editing software to integrate the new characters and voices into the original data.
[0090] Step 10:
[0091] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[0092] Step 11:
[0093] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[0094] Step 12:
[0095] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] In content such as videos and movies, it is desirable for users to replace characters with themselves to have a more personal and immersive experience. However, this process is technically complex, and there has been no system that makes it easy for users to do so. In addition, security measures to prevent inappropriate content generation are often lacking, so a means of safely providing customized content has been sought.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes means for transmitting the captured face and voice data to the server using a secure communication protocol, means for replacing the faces and voices of characters in the content with the face and voice of the user based on the input data, and means for changing the lines of other characters to new lines including the user's name, thereby enabling the user to safely and easily replace themselves with characters in the content and enjoy a personal and immersive experience.
[0101] "Means for capturing a user's face and voice in real time" refers to a system that uses a camera and a microphone to record and record a user's facial expressions and voice in real time.
[0102] "Means for inputting captured facial and audio data into generative AI" refers to the process or interface for inputting facial images and audio data captured in real time into an AI model.
[0103] "Means of replacing the faces and voices of characters in content with the face and voice of the user" refers to technology that converts and replaces the faces and voices of specific characters in content based on acquired face and voice data of the user.
[0104] The "means for changing the lines of other characters to new lines that include the user's name" is a text processing algorithm for modifying the lines spoken by other characters to include the user's name.
[0105] "Means for regenerating content" refers to methods or tools for editing and recomposing the original content based on the replaced face or voice data.
[0106] "Means for compressing and encoding the reproduced content" means techniques for reducing the size of the reproduced video and audio data and converting it into an appropriate format.
[0107] The "means for transmitting to the user's terminal" refers to a communication means for safely and quickly transmitting data generated from the server side to the user's device.
[0108] The "means for the user's terminal to play the regenerated content" refers to software or playback functions that allow the user's device to properly play the acquired customized content.
[0109] The present invention is a system that allows users to replace themselves with characters in movies and video content to enjoy a personalized viewing experience. The system includes a means for capturing the user's face and voice in real time and using a generative AI model to replace the face and voice of characters in the content with that of the user.
[0110] First, the user launches the application installed on their device, which can be a smartphone or tablet. The user grants camera and microphone access permissions from the application's main screen and captures their face and voice. This captured data is encoded in real time and sent to the server using a secure communication protocol (e.g., HTTPS).
[0111] The server uses a facial recognition algorithm (e.g., OpenCV) and a voice synthesis model (e.g., WaveNet) to analyze the received face and voice data. The server simultaneously retrieves the face, voice, movement, and dialogue data of the character from a database of the selected content. A generative AI model (e.g., DeepFake or DALL-E) is then used to replace the character's face and voice with the user's.
[0112] Additionally, a text processing algorithm is run to replace dialogue spoken by other characters with new dialogue that includes the user's name, causing other characters to call out the user's name. The server then uses this processed data to reconstruct the original video content in editing software such as Adobe Premiere Pro or Final Cut Pro. During this editing process, the new faces, voices, and dialogue are integrated into the timeline to create a seamless visual experience.
[0113] The regenerated content is compressed and encoded (e.g., H.264 codec) to optimize file size, and then transmitted to the device using a secure communication protocol. The user can then play the regenerated content stored in the device's local storage using the application's media player and enjoy customized movies and videos featuring themselves.
[0114] As a concrete example, suppose Person A captures his or her face and voice using an app and selects his or her favorite animated movie. The system inputs Person A's face and voice into a generative AI model, which then replaces the face and voice of the main character in the animated movie with Person A's. The dialogue of other characters in the anime is also changed to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie in which he or she appears as the main character.
[0115] Example prompt sentence:
[0116] "Mr. A, please face the camera to capture your face and speak into the microphone. Then, choose your favorite movie and press the start button."
[0117] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent inappropriate content from being generated.
[0118] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0119] Step 1:
[0120] Input: A user launches an application by tapping the app icon on the device's home screen.
[0121] Processing: The user launches the application by tapping the app icon on the device. The app displays a dialog requesting permission to use the camera and microphone.
[0122] Output: The application launches and displays a screen requesting camera and microphone access permissions.
[0123] What happens: The user taps the button to allow access permissions, granting access to the camera and microphone.
[0124] Step 2:
[0125] Input: The user faces the camera and speaks into the microphone for a few seconds.
[0126] Processing: The device captures the user's face and voice in real time and encodes it.
[0127] Output: The captured face and audio data is temporarily stored on the device.
[0128] What happens: The user looks at the camera and speaks into the microphone as instructed, and the application records and films this.
[0129] Step 3:
[0130] Input: Captured face and audio data.
[0131] Processing: The device sends the captured data to the server using a secure communication protocol (e.g., HTTPS).
[0132] Output: Face and voice data received by the server.
[0133] What it does: The application encrypts the captured data and sends it over the internet to a server.
[0134] Step 4:
[0135] Input: Face and voice data received by the server.
[0136] Processing: The server analyzes the data using facial recognition algorithms (e.g., OpenCV) and speech synthesis models (e.g., WaveNet).
[0137] Output: Face and voice data analyzed on the server.
[0138] Specific operation: The server analyzes facial features and voice waveforms to extract the necessary data.
[0139] Step 5:
[0140] Input: Analyzed face and audio data, selected content.
[0141] Processing: The server retrieves the character's face, voice, movement, and dialogue data from the database of the selected content.
[0142] Output: Character face data, voice data, movement data, and dialogue data.
[0143] What happens: The server queries the database for the specified content and retrieves the required data.
[0144] Step 6:
[0145] Input: Captured character data and parsed user data.
[0146] Processing: The server uses a generative AI model (e.g., DeepFake or DALL-E) to replace the character's face and voice with the user's face and voice.
[0147] Output: The replaced data.
[0148] Specific operation: The server runs the AI model and synthesizes the characters' faces and voices with those of the user.
[0149] Step 7:
[0150] Input: Replaced face and voice data and original description data.
[0151] What happens: The server changes the dialogue of the other characters to new dialogue that includes the user's name.
[0152] Output: The corrected dialogue data.
[0153] What it does: It uses a text processing algorithm to change the dialogue of characters to match the user's name.
[0154] Step 8:
[0155] Input: Replaced face and voice data and modified dialogue data.
[0156] Processing: The server uses editing software such as Adobe Premiere Pro or Final Cut Pro to reconstruct the content.
[0157] Output: The regenerated content.
[0158] How it works: The server synthesizes new data on the editing software's timeline to create a seamless video.
[0159] Step 9:
[0160] Input: The regenerated content.
[0161] Processing: The server compresses the regenerated content and encodes it with the H.264 codec.
[0162] Output: Compressed and encoded content data.
[0163] Specific operation: The server uses encoding software to optimize the file size and sends it to the device using a secure communication protocol.
[0164] Step 10:
[0165] Input: Compressed and encoded content data sent to the terminal.
[0166] Processing: The device saves the received content in local storage.
[0167] Output: Content saved to local storage.
[0168] Specific operation: The device receives the data and automatically saves it in the specified folder.
[0169] Step 11:
[0170] Input: Content stored in local storage.
[0171] Action: The device launches a media player and plays the stored content.
[0172] Output: The customized content that is played.
[0173] What happens: The user taps the play button in the app and starts watching the customized movie or video.
[0174] (Application example 1)
[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0176] Traditional movie and video content viewing experiences lack personalization and immersion for viewers. It is difficult for viewers to customize the experience by replacing characters with their own faces and voices, providing a limited experience. Furthermore, there are insufficient means to securely and efficiently transmit this customized content.
[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0178] In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into a generative AI model, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for regenerating the content based on the replaced data, means for transmitting the regenerated content to the user's terminal, means for the user to select content they wish to view, and means for saving the content in local storage and playing it back. This allows the user to securely and efficiently receive customized content in which they themselves are replaced with characters, and enjoy a highly immersive viewing experience.
[0179] "User" refers to the person who interacts with the system, captures their face and voice, and views customized content.
[0180] "Real-time face and voice capture method" refers to a device or software method for capturing a user's face and voice and collecting that data.
[0181] "Generative AI model" refers to an artificial intelligence model that uses captured facial and voice data to generate the face and voice of a new character.
[0182] "Input means" refers to the mechanism or method for inputting captured facial and voice data into a generative AI model.
[0183] "Characters" refers to characters that appear in movies and video content.
[0184] "Means of replacement" refers to the operation of using a generative AI model to replace the faces and voices of characters in the content with the user's face and voice.
[0185] "Regeneration means" refers to the method or process for creating new edited content based on replaced data.
[0186] "Means for playing content" refers to the mechanism or method for playing the created customized content on the user's terminal.
[0187] "Secure communication protocol" refers to a communication method for securely transmitting captured data to a server.
[0188] "Compression / Encoding Methods" means the methods and techniques that convert the data in the Reproduced Content into a format that can be efficiently transferred.
[0189] "Local storage" refers to the data storage area built into the user's device.
[0190] MODE FOR CARRYING OUT THE INVENTION
[0191] System program and processing overview
[0192] An embodiment of the present invention is a system that captures a user's face and voice to provide customized content. This system operates in cooperation with a terminal, such as a smartphone, smart glasses, a head-mounted display, or a robot, and a server. The specific operation of the system and its implementation method are described below.
[0193] Hardware and software configuration:
[0194] 1. Device:
[0195] Camera: Captures the user's face using OpenCV.
[0196] Microphone: Uses a SoundDevice to capture the user's voice.
[0197] Storage: Save data to local storage.
[0198] Communication: Use Requests to transmit data securely.
[0199] Media Player: Use MoviePy to play customized content.
[0200] 2. Server:
[0201] Generative AI model: Replaces characters in content based on the user's facial and voice data.
[0202] Database: Stores data on movies and video content.
[0203] Communication protocol: Use a communication protocol to securely transmit data.
[0204] Encoding software: compresses and encodes the reproduced content and sends it to the user's device.
[0205] Data processing and calculation flow:
[0206] 1. On the user's device:
[0207] The user launches an application on the device to capture their face and voice. The captured face image and voice data are temporarily stored in local storage.
[0208] The user selects the movie or video content they want to watch, and the content selection information is also stored in local storage.
[0209] 2. Server:
[0210] The facial and voice data sent from the device is received and input into the generative AI model, and this process uses a secure communication protocol to protect the data.
[0211] The generative AI model replaces the faces and voices of characters in the content with those of the user based on the input data.
[0212] Once all processing is complete, the regenerated content is compressed and encoded using encoding software and sent to the user's device.
[0213] 3. Playing content:
[0214] The user's device receives the new content and saves it to local storage, where the user can play the saved customized content using MoviePy.
[0215] Examples:
[0216] For example, if a user captures their face and voice and selects a particular movie, they can send the data to a server using a prompt like this:
[0217] bash
[0218] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[0219] This prompt example uploads the files face.jpg and audio.wav to the server and specifies the selected movie ID (example_movie_id). The server generates customized content based on the received data and sends it to the user's device.
[0220] As described above, by using this system, users can have a personalized experience by replacing themselves with characters in movies and video content based on their facial and voice data.
[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0222] Step 1:
[0223] The user launches an application on the device. The user opens the application by tapping the app icon on the device's home screen. The user then proceeds to the face and voice capture screen.
[0224] Input: Face and voice capture request
[0225] Output: Face image (face.jpg) and audio data (audio.wav)
[0226] Specific behavior:
[0227] The device will request permission to access the camera and microphone and display a face and audio capture page. When the user faces the camera and speaks into the microphone, the device will capture video and audio in real time and save them to local storage as "face.jpg" and "audio.wav."
[0228] Step 2:
[0229] The user selects the movie or video content they want to watch by using the application's interface to select the title from the content library.
[0230] Input: User-selected content ID
[0231] Output: Selected content ID (content_id)
[0232] Specific behavior:
[0233] The device interface displays a content selection screen to the user, and the user selects the content they want to watch by tapping on it. This selection information is saved as "content_id."
[0234] Step 3:
[0235] The device transmits the captured face and voice data and the selected content ID to a server using a secure communication protocol.
[0236] Input: Face image (face.jpg), audio data (audio.wav), content ID (content_id)
[0237] Output: Send data to the server
[0238] Specific behavior:
[0239] The terminal uses the Requests library to send data to the server using a secure communication protocol, specifically using the following prompt sentence:
[0240] bash
[0241] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[0242] Step 4:
[0243] The server analyzes the received facial and voice data and inputs it into a generative AI model, which then replaces the face and voice of the characters in the content with those of the user based on their face and voice.
[0244] Input: Face image (face.jpg), audio data (audio.wav)
[0245] Output: Replaced face and voice data
[0246] Specific behavior:
[0247] The server uses a generative AI model to analyze the user's facial image and voice data and process it to replace them with a character in the content, converting the character's face and voice to that of the user.
[0248] Step 5:
[0249] The server regenerates the content based on the replaced face and voice data, compresses and encodes the regenerated content, and sends it to the user's device.
[0250] Input: Replaced face and voice data
[0251] Output: Regenerated content files
[0252] Specific behavior:
[0253] The server uses encoding software to compress and encode the reproduced content into an efficiently transferable format, and the encoded content file is sent to the user's device using a secure communications protocol.
[0254] Step 6:
[0255] The device receives the regenerated content and stores it in local storage, allowing the user to play and enjoy the content.
[0256] Input: Regenerated content file
[0257] Output: Playable content file
[0258] Specific behavior:
[0259] The device receives the regenerated content sent from the server and stores it in local storage. The user can then use MoviePy to play the saved customized content and enjoy watching it.
[0260] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0261] The present invention provides a more interactive and emotional experience by combining a system for personalizing the viewing experience of movies and video content with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0262] User operations
[0263] 1. Launching the application
[0264] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[0265] 2. Face and voice capture
[0266] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0267] When a user faces the camera and speaks into the microphone, the device captures video and audio in real time and records them as digital data.
[0268] 3. Analysis by Emotion Engine
[0269] The captured facial and voice data is input into the emotion engine, which analyzes the user's emotions. The device sends the user's facial expressions and tone of voice to the emotion engine, which then generates emotion data.
[0270] 4. Content Selection
[0271] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[0272] Server Processing
[0273] 1. Data Receipt and Analysis
[0274] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[0275] 2. Acquiring character data
[0276] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0277] 3. Use of generative AI
[0278] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[0279] 4. Generating new lines
[0280] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0281] 5. Building regenerative content
[0282] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[0283] 6. Data Compression and Transfer
[0284] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[0285] Terminal handling
[0286] 1. Receiving new content
[0287] The terminal receives the regenerated content sent from the server and stores it in local storage.
[0288] 2. Playing content
[0289] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0290] Specific examples
[0291] For example, suppose Person B captures his or her face and voice, and the emotion engine detects "joy." If Person B selects an action movie, the server uses Person B's emotional data to adjust the facial expression and tone of voice of the protagonist in the action scene to match joy. Furthermore, the server changes the dialogue of other characters to call Person B's name, generating a new action movie that reflects the emotional expression. This new content is compressed, encoded, and sent to Person B's device. Person B can then play this customized action movie in the application and enjoy a more emotionally resonant viewing experience.
[0292] This system not only allows users to have a more interactive and emotional experience, but also provides security to prevent the generation of inappropriate content.
[0293] The processing flow will be explained below.
[0294] Step 1:
[0295] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[0296] Step 2:
[0297] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0298] Step 3:
[0299] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[0300] Step 4:
[0301] The device sends the captured data to the emotion engine, which analyzes the user's facial expressions and tone of voice to generate emotion data.
[0302] Step 5:
[0303] The device displays a content selection interface, allowing the user to select the movie or video content they want to view.
[0304] Step 6:
[0305] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[0306] Step 7:
[0307] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0308] Step 8:
[0309] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[0310] Step 9:
[0311] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0312] Step 10:
[0313] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[0314] Step 11:
[0315] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[0316] Step 12:
[0317] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[0318] Step 13:
[0319] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0320] Example 2
[0321] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0322] Conventional movie and video content has primarily been a one-way viewing experience, making it difficult to customize it to reflect the user's emotions and individual characteristics. The generation of content that interactively responds to the viewer's emotions has been extremely limited. Furthermore, conventional technology has been inadequate for reflecting the user's face and voice onto characters in movies and videos in real time. By solving these issues, it was necessary to provide a more personal and emotional viewing experience.
[0323] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing the facial expressions and movements of characters in detail from a database, a means for analyzing the line database to obtain line lists of other characters and generate new lines, and a means for replacing the face and voice of a character based on the face and voice data of a user using the generated AI model. This makes it possible to replace the face and voice of a character in selected content with the face and voice of the user, thereby interactively customizing the content according to the user's emotions.
[0324] "User" means an individual who utilizes the System to customize movie and video content.
[0325] "Device" means a device that uses a camera and microphone to capture face and voice, including a smartphone, tablet, or computer.
[0326] "Capture" refers to the process of using a camera and microphone to digitally record facial and audio data.
[0327] "Emotion engine" refers to software that analyzes a user's emotions from captured facial and voice data and generates emotional data.
[0328] "Generative AI model" refers to an artificial intelligence model that replaces the faces and voices of characters in movies and video content with those of the user based on the user's facial and voice data.
[0329] A "dialogue database" refers to a database that stores dialogue used in movies and video content and keeps it in an analyzable format.
[0330] "Secure communication protocol" refers to the communication protocol for securely transmitting captured data to a server.
[0331] "Regenerated Content" refers to customized film or video content generated based on a user's face and voice, emotional data, and new dialogue data.
[0332] "Editing Software" means video editing tools used to create Regenerated Content, including, but not limited to, Adobe Premiere Pro.
[0333] "Compression / encoding" refers to the process of converting digital content to a specific bit rate and format to reduce the data size.
[0334] "Local storage" refers to a memory area that exists within a device and is used to store data.
[0335] The present invention provides a system that provides a more interactive and emotional experience by customizing the viewing experience of movies and video content and combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[0336] Hardware and software used
[0337] This system uses devices such as smartphones, tablets, and PCs. Each device must be equipped with a camera and microphone. The server is a back-end server that receives data using a secure communication protocol (e.g., HTTPS).
[0338] Additionally, the following specific software is used:
[0339] Emotion engine: Emotion recognition algorithm (e.g. Microsoft Azure Emotion API)
[0340] Generative AI models: Generative AI models (e.g., OpenAI GPT-3, DeepFaceLab)
[0341] Editing software: Video editing software (e.g. Adobe Premiere Pro)
[0342] Specific implementation methods
[0343] 1. Launching the application
[0344] The user launches the application by tapping the app icon on the device's home screen, and the application displays its initial screen.
[0345] 2. Face and voice capture
[0346] 1. The device will request permission to access the camera and microphone. If the user allows it, the face and voice capture page will open.
[0347] 2. The user faces the camera and speaks as instructed. The device captures this in real time and generates digital data.
[0348] 3. Analysis by Emotion Engine
[0349] The device sends the captured face and voice data to the emotion engine, which analyzes the data and generates emotion data.
[0350] 4. Content Selection
[0351] The user selects their favorite movie from the "Choose a Movie" interface. Once the user makes a selection, the device displays detailed information from the content library.
[0352] 5. Data processing by the server
[0353] 1. The device sends the capture data and emotion data to the server, which then retrieves the character data for the selected movie from the database.
[0354] 2. The server uses a generative AI model to replace the character based on the user's facial and voice data, adjusting facial expressions and tone of voice based on emotional data.
[0355] 3. The server analyzes the dialogue database and generates new dialogue.
[0356] 4. The server uses editing software such as Adobe Premiere Pro to integrate the new characters, voices, and emotional expressions into the original data.
[0357] 6. Content Compression and Transfer
[0358] The server compresses the regenerated content using an appropriate video codec, such as H.264, and sends it to the device.
[0359] 7. Receiving and Playing Content
[0360] 1. The device receives the regenerated content from the server and stores it in local storage.
[0361] 2. The user selects new content within the application and begins playback.
[0362] Specific examples
[0363] For example, if a user selects an action movie and captures their face and voice, the emotion engine may recognize the user's emotion as "joy." In this case, the server will replace the main character of the movie with the user's own character based on the user's face and voice data, and adjust the characters' facial expressions and tone of voice to reflect "joy." It will also change the dialogue so that other characters call the user's name, creating a customized action movie. This new content is then sent to the user's device, providing the user with an interactive and emotional viewing experience.
[0364] Prompt Sentence Examples
[0365] Example prompts for generative AI models:
[0366] "Based on the user's facial and voice data, replace the face and voice of the main character in an action movie with the user's. Also, if the emotion data is 'joy', adjust the main character's facial expression and tone of voice to be happy, and generate new lines for other characters in the movie to call the user's name."
[0367] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0368] Step 1:
[0369] The user launches an application on the device.
[0370] Specific operation: The user taps the app icon on the device's home screen and the application launches.
[0371] Input: User taps.
[0372] Output: The initial screen of the application is displayed.
[0373] Step 2:
[0374] The device will request permission to access the camera and microphone.
[0375] Specific behavior: The application displays "Please allow access to your camera and microphone." The user taps "Allow."
[0376] Input: User's access permission permission operation.
[0377] Output: The camera and microphone will be activated and the Face and Voice Capture page will open.
[0378] Step 3:
[0379] The user faces the camera and speaks into the microphone, which the device captures as digital data in real time.
[0380] Specific operation: The user follows the instructions to face the camera and say "hello." The device captures the facial image and audio and records them as digital data.
[0381] Input: User's face and voice.
[0382] Output: Digital data of the captured face and voice.
[0383] Step 4:
[0384] The device sends the captured face and voice data to the emotion engine.
[0385] How it works: The device uploads face and voice data to the cloud-based emotion engine, which then analyzes the data and generates emotion data.
[0386] Input: Captured digital face and voice data.
[0387] Output: Parsed emotion data.
[0388] Step 5:
[0389] The user selects the movie or video content they want to view on the device interface.
[0390] What happens: The user selects a movie from the content library and confirms the selection.
[0391] Input: User's movie title selection.
[0392] Output: Detailed information about the selected movie is displayed.
[0393] Step 6:
[0394] The device transmits the capture data and emotion data to the server.
[0395] Specific operation: The device encrypts the data and sends it to the server via a secure communication protocol.
[0396] Input: Capture data, emotion data.
[0397] Output: The server receives face, voice, and emotion data.
[0398] Step 7:
[0399] The server retrieves character data from the database and analyzes it.
[0400] Specific operation: The server accesses a movie database to obtain data on the characters' faces, voices, movements, and lines. The obtained data is then analyzed using a facial recognition algorithm and a voice synthesis model.
[0401] Input: A content database containing character data.
[0402] Output: Analyzed character face, voice, movement, and dialogue data.
[0403] Step 8:
[0404] The server uses a generative AI model to replace characters based on the user's facial and voice data.
[0405] How it works: The server uses a generative AI model to replace the user's face and voice data with the character's face and voice, and also adjusts facial expressions and tone of voice based on emotional data.
[0406] Input: Capture data, emotion data, character data.
[0407] Output: Character data replaced with the user's face and voice.
[0408] Step 9:
[0409] The server analyzes the dialogue database and generates new dialogue.
[0410] How it works: The server accesses the dialogue database and generates new dialogue for other characters to call the user's name. The dialogue generation uses a generative AI model.
[0411] Input: Dialogue database, user's name.
[0412] Output: New dialogue data containing the user's name.
[0413] Step 10:
[0414] The server recreates the original video content.
[0415] Specific operation: The server uses editing software such as Adobe Premiere Pro to integrate the generated face, voice, emotion data, and new dialogue data into the original video and re-edit it.
[0416] Input: Generated face, voice, and emotion data, new dialogue data, and original video content.
[0417] Output: The regenerated video content.
[0418] Step 11:
[0419] The server compresses and encodes the regenerated content and sends it to the device.
[0420] Specific operation: The server compresses the reproduced content using a video codec such as H.264 and sends it to the terminal via a secure communication protocol.
[0421] Input: The regenerated content.
[0422] Output: Compressed and encoded regenerated content.
[0423] Step 12:
[0424] The device receives the regenerated content and stores it in local storage.
[0425] Specific operation: The device receives the regenerated content from the server through a secure communication protocol and stores it in local storage.
[0426] Input: Compressed and encoded regenerated content.
[0427] Output: Regenerated content saved to local storage.
[0428] Step 13:
[0429] The user starts playing the regenerated content within the application on the terminal.
[0430] Specific operation: The user operates the application, selects the regenerated content, and presses the play button. The device's media player starts and the regenerated content is played.
[0431] Input: User playback operations.
[0432] Output: Playback of regenerated content.
[0433] The above processing steps provide the user with a customized and interactive viewing experience.
[0434] (Application example 2)
[0435] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0436] Traditional movie and video content could only provide fixed content and could not be customized according to the viewer's emotions or individual characteristics. This made it difficult to provide a more interactive and emotional viewing experience. Furthermore, it was not possible to reflect the viewer's name or individual characteristics, which resulted in an insufficient personalized experience.
[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into an emotion engine and analyzing it, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for adjusting the character's facial expression and tone of voice based on emotion data, means for changing the lines of other characters to new lines including the user's name, and means for regenerating content based on the replaced data, the adjusted data, and the changed lines and transmitting the regenerated content to the user's terminal. This makes it possible to provide an interactive and emotional viewing experience customized for each user.
[0438] "Means for capturing the user's face and voice in real time" refers to devices or software that instantly record the facial image and voice of the user as they speak to the terminal.
[0439] "Emotion engine" refers to algorithms and software that analyze captured facial and voice data to infer a user's emotional state.
[0440] "Means for replacing the face and voice of a character in content with the user's face and voice based on input data" refers to devices or software that execute a process to replace the face and voice of a character appearing in content using collected data of the user's face and voice.
[0441] "Means for adjusting the facial expressions and tone of voice of characters" refers to devices or software that change the facial expressions and voice characteristics of characters in the content based on emotional data obtained by the emotion engine.
[0442] "Means for changing the lines of other characters into new lines that include the user's name" refers to devices or software that convert the words spoken by characters included in the content into new words that include the user's name.
[0443] "Means for regenerating content" refers to devices or software that regenerate new video and audio data based on the user's face, voice, emotional data, and changed lines.
[0444] "Means for transmitting the regenerated content to the user's device" means the equipment or software that transmits the edited or generated new video and audio to the user's device using the appropriate protocol.
[0445] This invention is a system for customizing a user's viewing experience, implemented using a smartphone application. Specifically, it captures the user's face and voice, analyzes the data with an emotion engine, and uses a generative AI model to customize the faces, voices, emotions, and lines of characters in video content in real time, providing an interactive and emotional viewing experience.
[0446] Hardware and Software Use
[0447] Device: A smartphone is used. The smartphone's front camera and microphone are used to capture face and voice. The smartphone must also have a media player installed for viewing.
[0448] Server: Runs the emotion engine, generative AI model, and dialogue conversion algorithm.
[0449] Emotion engine: Analyzes captured facial and voice data to interpret user emotions in real time. Specific software used includes facial recognition algorithms (e.g., OpenCV) and speech recognition software (e.g., the speech_recognition library).
[0450] Generative AI models: Generate the faces, voices, and expressions of characters in content based on emotional and captured data. Specific examples include face-swapping algorithms and voice synthesis models (e.g., DeepFake, Tacotron).
[0451] Dialogue conversion algorithm: An algorithm to change the dialogue of other characters into new dialogue that includes the user's name. New dialogue is generated using a natural language processing model (e.g., GPT-3).
[0452] System flow
[0453] 1. User operations
[0454] A user launches a smartphone application and uses the camera and microphone to capture their face and voice.
[0455] 2. Data Analysis
[0456] The captured face and voice data is sent to the emotion engine, which analyzes the user's emotional state and formats the results as emotion data, including the user's facial expressions and tone of voice.
[0457] 3. Server Processing
[0458] The server receives the emotion data and inputs it into a generative AI model. The face and voice of the characters in the content are replaced with the user's, and the facial expressions and tone of voice are adjusted. The lines of other characters are also converted into new lines that include the user's name.
[0459] 4. Regenerating and Submitting Content
[0460] The server regenerates the video content based on the adjusted data and the new dialogue. This new content is compressed, encoded, and sent to the user's smartphone using a secure communication protocol.
[0461] 5. Playing Content
[0462] The user's smartphone receives the regenerated content and launches a media player to play it.
[0463] Adding specific examples
[0464] For example, a user opens the app and captures their face and voice. The emotion engine then detects "happiness." Based on this result, the facial expression and voice of the protagonist of the action movie selected by the user are adjusted to reflect a state of happiness. Furthermore, the dialogue of other characters is changed to call the user's name, "Takashi."
[0465] Example prompts for generative AI models
[0466] text
[0467] "User emotion analysis data": "Joy",
[0468] "Character Layer": {
[0469] "Expression": "Joy",
[0470] "Tone": "bright",
[0471] "Name": "Takashi"
[0472] },
[0473] "Dialogue Database": "Generate new dialogue"
[0474] The customized content thus generated is then sent to the user's smartphone, allowing the user to enjoy a more emotionally resonant viewing experience.
[0475] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0476] Step 1:
[0477] A user launches a smartphone application. The user taps the app icon on the device's home screen to open it. The application requests permission to access the camera and microphone and displays a face and voice capture page. The input is the user's operation, and the output is the activation of the camera and microphone.
[0478] Step 2:
[0479] The user faces the camera and speaks into the microphone. The device captures video and records audio in real time. The captured data is temporarily stored in the device's local storage. The input is the user's face and voice, and the output is facial image data and audio data.
[0480] Step 3:
[0481] The device sends the captured facial image data and voice data to the emotion engine to analyze the emotional state. The emotion engine uses facial recognition algorithms and voice analysis algorithms (e.g., OpenCV, speech_recognition) to analyze the user's facial expressions and tone of voice. The input is facial image data and voice data, and the output is the user's emotional data (e.g., "happiness," "excitement," etc.).
[0482] Step 4:
[0483] The device generates emotion data and transmits it to the server along with the captured data using a secure communication protocol (e.g., HTTPS). The input is facial image data, voice data, and emotion data, and the output is the securely transmitted data.
[0484] Step 5:
[0485] The server inputs the received facial image data, voice data, and emotion data into a generative AI model. The generative AI model then replaces the faces and voices of characters in the content with the user's face and voice. The input is the captured data and emotion data, and the output is the converted face and voice of the characters.
[0486] Step 6:
[0487] The server adjusts the facial expressions and tone of voice of the characters in the content based on the emotional data. The generative AI model changes the facial expressions and voice of the characters in real time to match the user's emotions. The input is emotional data, and the output is adjusted facial expression data and tone of voice data.
[0488] Step 7:
[0489] The server analyzes the dialogue of other characters and converts it into new dialogue that includes the user's name. It generates the dialogue using a natural language processing algorithm (e.g., GPT-3). The input is the existing dialogue data, and the output is the new dialogue data.
[0490] Step 8:
[0491] The server regenerates the content based on the facial data, voice data, adjusted facial expression data, and new dialogue data. This data is then integrated using video editing software (e.g., Adobe Premiere Pro). The input is all the converted data, and the output is the regenerated content.
[0492] Step 9:
[0493] The server compresses and encodes the regenerated content, encoding it at the appropriate bitrate, and sends the compressed data to the device using a secure communication protocol. The input is the regenerated content, and the output is the compressed and encoded content.
[0494] Step 10:
[0495] The device receives the regenerated content and saves it to local storage. It then launches the device's media player and begins playing the new content. The input is the compressed and encoded content, and the output is a customized video played to the user.
[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0497] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0498] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0499] [Second embodiment]
[0500] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0501] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0502] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0504] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0506] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0507] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0510] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0511] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0512] The present invention is a system for personalizing the viewing experience of movies and video content. Users can use their own face and voice to replace characters in the content with themselves. This system is characterized by exchanging data using a secure communication protocol and playing the generated customized content on the user's terminal. Specific embodiments of the system are described below.
[0513] User operations
[0514] 1. Launching the application
[0515] The user launches an application on the device by tapping the app icon on the device's home screen.
[0516] 2. Face and voice capture
[0517] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0518] When a user faces the camera and speaks into the microphone for a few seconds, the device captures and records video and audio in real time.
[0519] 3. Content Selection
[0520] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[0521] Server Processing
[0522] 1. Data Receipt and Analysis
[0523] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[0524] 2. Acquiring character data
[0525] The server retrieves character data (face, voice, movements, lines) from the database of the selected content.
[0526] 3. Use of generative AI
[0527] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the inputted face and voice data of the user.
[0528] 4. Generating new lines
[0529] The server changes the lines that other characters use to call out to the user to new lines that include the user's name, so that other characters will call out the user's name.
[0530] 5. Building regenerative content
[0531] The server reconstructs the original video content based on the generated facial, voice, and dialogue data, and then uses editing software to integrate the new characters and voices into the original data.
[0532] 6. Data Compression and Transfer
[0533] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[0534] Terminal handling
[0535] 1. Receiving new content
[0536] The terminal receives the regenerated content sent from the server and stores it in local storage.
[0537] 2. Playing content
[0538] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0539] Specific examples
[0540] For example, suppose Person A captures his or her own face and voice and selects an animated movie. The server inputs Person A's face and voice into the generation AI, which then replaces the face and voice of the main character in the animated movie with Person A's. It also changes the lines of other characters to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie starring himself or herself as the main character.
[0541] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent the generation of inappropriate content.
[0542] The processing flow will be explained below.
[0543] Step 1:
[0544] The user launches an application on the device by tapping the corresponding app icon on the device's home screen.
[0545] Step 2:
[0546] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0547] Step 3:
[0548] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[0549] Step 4:
[0550] The user selects the movie or video content they want to watch through the device's interface, and the device selects the title from its content library.
[0551] Step 5:
[0552] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[0553] Step 6:
[0554] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0555] Step 7:
[0556] The server uses generative AI to replace the face and voice of the characters in the content with those of the user based on the inputted face and voice data of the user, and generates a digital model that reflects the user's characteristics.
[0557] Step 8:
[0558] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0559] Step 9:
[0560] The server recreates the original video content based on the generated facial, voice, and dialogue data, and uses editing software to integrate the new characters and voices into the original data.
[0561] Step 10:
[0562] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[0563] Step 11:
[0564] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[0565] Step 12:
[0566] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0567] Example 1
[0568] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0569] In content such as videos and movies, it is desirable for users to replace characters with themselves to have a more personal and immersive experience. However, this process is technically complex, and there has been no system that makes it easy for users to do so. In addition, security measures to prevent inappropriate content generation are often lacking, so a means of safely providing customized content has been sought.
[0570] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0571] In this invention, the server includes means for transmitting the captured face and voice data to the server using a secure communication protocol, means for replacing the faces and voices of characters in the content with the face and voice of the user based on the input data, and means for changing the lines of other characters to new lines including the user's name, thereby enabling the user to safely and easily replace themselves with characters in the content and enjoy a personal and immersive experience.
[0572] "Means for capturing a user's face and voice in real time" refers to a system that uses a camera and a microphone to record and record a user's facial expressions and voice in real time.
[0573] "Means for inputting captured facial and audio data into generative AI" refers to the process or interface for inputting facial images and audio data captured in real time into an AI model.
[0574] "Means of replacing the faces and voices of characters in content with the face and voice of the user" refers to technology that converts and replaces the faces and voices of specific characters in content based on acquired face and voice data of the user.
[0575] The "means for changing the lines of other characters to new lines that include the user's name" is a text processing algorithm for modifying the lines spoken by other characters to include the user's name.
[0576] "Means for regenerating content" refers to methods or tools for editing and recomposing the original content based on the replaced face or voice data.
[0577] "Means for compressing and encoding the reproduced content" means techniques for reducing the size of the reproduced video and audio data and converting it into an appropriate format.
[0578] The "means for transmitting to the user's terminal" refers to a communication means for safely and quickly transmitting data generated from the server side to the user's device.
[0579] The "means for the user's terminal to play the regenerated content" refers to software or playback functions that allow the user's device to properly play the acquired customized content.
[0580] The present invention is a system that allows users to replace themselves with characters in movies and video content to enjoy a personalized viewing experience. The system includes a means for capturing the user's face and voice in real time and using a generative AI model to replace the face and voice of characters in the content with that of the user.
[0581] First, the user launches the application installed on their device, which can be a smartphone or tablet. The user grants camera and microphone access permissions from the application's main screen and captures their face and voice. This captured data is encoded in real time and sent to the server using a secure communication protocol (e.g., HTTPS).
[0582] The server uses a facial recognition algorithm (e.g., OpenCV) and a voice synthesis model (e.g., WaveNet) to analyze the received face and voice data. The server simultaneously retrieves the face, voice, movement, and dialogue data of the character from a database of the selected content. A generative AI model (e.g., DeepFake or DALL-E) is then used to replace the character's face and voice with the user's.
[0583] Additionally, a text processing algorithm is run to replace dialogue spoken by other characters with new dialogue that includes the user's name, causing other characters to call out the user's name. The server then uses this processed data to reconstruct the original video content in editing software such as Adobe Premiere Pro or Final Cut Pro. During this editing process, the new faces, voices, and dialogue are integrated into the timeline to create a seamless visual experience.
[0584] The regenerated content is compressed and encoded (e.g., H.264 codec) to optimize file size, and then transmitted to the device using a secure communication protocol. The user can then play the regenerated content stored in the device's local storage using the application's media player and enjoy customized movies and videos featuring themselves.
[0585] As a concrete example, suppose Person A captures his or her face and voice using an app and selects his or her favorite animated movie. The system inputs Person A's face and voice into a generative AI model, which then replaces the face and voice of the main character in the animated movie with Person A's. The dialogue of other characters in the anime is also changed to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie in which he or she appears as the main character.
[0586] Example prompt sentence:
[0587] "Mr. A, please face the camera to capture your face and speak into the microphone. Then, choose your favorite movie and press the start button."
[0588] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent inappropriate content from being generated.
[0589] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0590] Step 1:
[0591] Input: A user launches an application by tapping the app icon on the device's home screen.
[0592] Processing: The user launches the application by tapping the app icon on the device. The app displays a dialog requesting permission to use the camera and microphone.
[0593] Output: The application launches and displays a screen requesting camera and microphone access permissions.
[0594] What happens: The user taps the button to allow access permissions, granting access to the camera and microphone.
[0595] Step 2:
[0596] Input: The user faces the camera and speaks into the microphone for a few seconds.
[0597] Processing: The device captures the user's face and voice in real time and encodes it.
[0598] Output: The captured face and audio data is temporarily stored on the device.
[0599] What happens: The user looks at the camera and speaks into the microphone as instructed, and the application records and films this.
[0600] Step 3:
[0601] Input: Captured face and audio data.
[0602] Processing: The device sends the captured data to the server using a secure communication protocol (e.g., HTTPS).
[0603] Output: Face and voice data received by the server.
[0604] What it does: The application encrypts the captured data and sends it over the internet to a server.
[0605] Step 4:
[0606] Input: Face and voice data received by the server.
[0607] Processing: The server analyzes the data using facial recognition algorithms (e.g., OpenCV) and speech synthesis models (e.g., WaveNet).
[0608] Output: Face and voice data analyzed on the server.
[0609] Specific operation: The server analyzes facial features and voice waveforms to extract the necessary data.
[0610] Step 5:
[0611] Input: Analyzed face and audio data, selected content.
[0612] Processing: The server retrieves the character's face, voice, movement, and dialogue data from the database of the selected content.
[0613] Output: Character face data, voice data, movement data, and dialogue data.
[0614] What happens: The server queries the database for the specified content and retrieves the required data.
[0615] Step 6:
[0616] Input: Captured character data and parsed user data.
[0617] Processing: The server uses a generative AI model (e.g., DeepFake or DALL-E) to replace the character's face and voice with the user's face and voice.
[0618] Output: The replaced data.
[0619] Specific operation: The server runs the AI model and synthesizes the characters' faces and voices with those of the user.
[0620] Step 7:
[0621] Input: Replaced face and voice data and original description data.
[0622] What happens: The server changes the dialogue of the other characters to new dialogue that includes the user's name.
[0623] Output: The corrected dialogue data.
[0624] What it does: It uses a text processing algorithm to change the dialogue of characters to match the user's name.
[0625] Step 8:
[0626] Input: Replaced face and voice data and modified dialogue data.
[0627] Processing: The server uses editing software such as Adobe Premiere Pro or Final Cut Pro to reconstruct the content.
[0628] Output: The regenerated content.
[0629] How it works: The server synthesizes new data on the editing software's timeline to create a seamless video.
[0630] Step 9:
[0631] Input: The regenerated content.
[0632] Processing: The server compresses the regenerated content and encodes it with the H.264 codec.
[0633] Output: Compressed and encoded content data.
[0634] Specific operation: The server uses encoding software to optimize the file size and sends it to the device using a secure communication protocol.
[0635] Step 10:
[0636] Input: Compressed and encoded content data sent to the terminal.
[0637] Processing: The device saves the received content in local storage.
[0638] Output: Content saved to local storage.
[0639] Specific operation: The device receives the data and automatically saves it in the specified folder.
[0640] Step 11:
[0641] Input: Content stored in local storage.
[0642] Action: The device launches a media player and plays the stored content.
[0643] Output: The customized content that is played.
[0644] What happens: The user taps the play button in the app and starts watching the customized movie or video.
[0645] (Application example 1)
[0646] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0647] Traditional movie and video content viewing experiences lack personalization and immersion for viewers. It is difficult for viewers to customize the experience by replacing characters with their own faces and voices, providing a limited experience. Furthermore, there are insufficient means to securely and efficiently transmit this customized content.
[0648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0649] In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into a generative AI model, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for regenerating the content based on the replaced data, means for transmitting the regenerated content to the user's terminal, means for the user to select content they wish to view, and means for saving the content in local storage and playing it back. This allows the user to securely and efficiently receive customized content in which they themselves are replaced with characters, and enjoy a highly immersive viewing experience.
[0650] "User" refers to the person who interacts with the system, captures their face and voice, and views customized content.
[0651] "Real-time face and voice capture method" refers to a device or software method for capturing a user's face and voice and collecting that data.
[0652] "Generative AI model" refers to an artificial intelligence model that uses captured facial and voice data to generate the face and voice of a new character.
[0653] "Input means" refers to the mechanism or method for inputting captured facial and voice data into a generative AI model.
[0654] "Characters" refers to characters that appear in movies and video content.
[0655] "Means of replacement" refers to the operation of using a generative AI model to replace the faces and voices of characters in the content with the user's face and voice.
[0656] "Regeneration means" refers to the method or process for creating new edited content based on replaced data.
[0657] "Means for playing content" refers to the mechanism or method for playing the created customized content on the user's terminal.
[0658] "Secure communication protocol" refers to a communication method for securely transmitting captured data to a server.
[0659] "Compression / Encoding Methods" means the methods and techniques that convert the data in the Reproduced Content into a format that can be efficiently transferred.
[0660] "Local storage" refers to the data storage area built into the user's device.
[0661] MODE FOR CARRYING OUT THE INVENTION
[0662] System program and processing overview
[0663] An embodiment of the present invention is a system that captures a user's face and voice to provide customized content. This system operates in cooperation with a terminal, such as a smartphone, smart glasses, a head-mounted display, or a robot, and a server. The specific operation of the system and its implementation method are described below.
[0664] Hardware and software configuration:
[0665] 1. Device:
[0666] Camera: Captures the user's face using OpenCV.
[0667] Microphone: Uses a SoundDevice to capture the user's voice.
[0668] Storage: Save data to local storage.
[0669] Communication: Use Requests to transmit data securely.
[0670] Media Player: Use MoviePy to play customized content.
[0671] 2. Server:
[0672] Generative AI model: Replaces characters in content based on the user's facial and voice data.
[0673] Database: Stores data on movies and video content.
[0674] Communication protocol: Use a communication protocol to securely transmit data.
[0675] Encoding software: compresses and encodes the reproduced content and sends it to the user's device.
[0676] Data processing and calculation flow:
[0677] 1. On the user's device:
[0678] The user launches an application on the device to capture their face and voice. The captured face image and voice data are temporarily stored in local storage.
[0679] The user selects the movie or video content they want to watch, and the content selection information is also stored in local storage.
[0680] 2. Server:
[0681] The facial and voice data sent from the device is received and input into the generative AI model, and this process uses a secure communication protocol to protect the data.
[0682] The generative AI model replaces the faces and voices of characters in the content with those of the user based on the input data.
[0683] Once all processing is complete, the regenerated content is compressed and encoded using encoding software and sent to the user's device.
[0684] 3. Playing content:
[0685] The user's device receives the new content and saves it to local storage, where the user can play the saved customized content using MoviePy.
[0686] Examples:
[0687] For example, if a user captures their face and voice and selects a particular movie, they can send the data to a server using a prompt like this:
[0688] bash
[0689] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[0690] This prompt example uploads the files face.jpg and audio.wav to the server and specifies the selected movie ID (example_movie_id). The server generates customized content based on the received data and sends it to the user's device.
[0691] As described above, by using this system, users can have a personalized experience by replacing themselves with characters in movies and video content based on their facial and voice data.
[0692] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0693] Step 1:
[0694] The user launches an application on the device. The user opens the application by tapping the app icon on the device's home screen. The user then proceeds to the face and voice capture screen.
[0695] Input: Face and voice capture request
[0696] Output: Face image (face.jpg) and audio data (audio.wav)
[0697] Specific behavior:
[0698] The device will request permission to access the camera and microphone and display a face and audio capture page. When the user faces the camera and speaks into the microphone, the device will capture video and audio in real time and save them to local storage as "face.jpg" and "audio.wav."
[0699] Step 2:
[0700] The user selects the movie or video content they want to watch by using the application's interface to select the title from the content library.
[0701] Input: User-selected content ID
[0702] Output: Selected content ID (content_id)
[0703] Specific behavior:
[0704] The device interface displays a content selection screen to the user, and the user selects the content they want to watch by tapping on it. This selection information is saved as "content_id."
[0705] Step 3:
[0706] The device transmits the captured face and voice data and the selected content ID to a server using a secure communication protocol.
[0707] Input: Face image (face.jpg), audio data (audio.wav), content ID (content_id)
[0708] Output: Send data to the server
[0709] Specific behavior:
[0710] The terminal uses the Requests library to send data to the server using a secure communication protocol, specifically using the following prompt sentence:
[0711] bash
[0712] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[0713] Step 4:
[0714] The server analyzes the received facial and voice data and inputs it into a generative AI model, which then replaces the face and voice of the characters in the content with those of the user based on their face and voice.
[0715] Input: Face image (face.jpg), audio data (audio.wav)
[0716] Output: Replaced face and voice data
[0717] Specific behavior:
[0718] The server uses a generative AI model to analyze the user's facial image and voice data and process it to replace them with a character in the content, converting the character's face and voice to that of the user.
[0719] Step 5:
[0720] The server regenerates the content based on the replaced face and voice data, compresses and encodes the regenerated content, and sends it to the user's device.
[0721] Input: Replaced face and voice data
[0722] Output: Regenerated content files
[0723] Specific behavior:
[0724] The server uses encoding software to compress and encode the reproduced content into an efficiently transferable format, and the encoded content file is sent to the user's device using a secure communications protocol.
[0725] Step 6:
[0726] The device receives the regenerated content and stores it in local storage, allowing the user to play and enjoy the content.
[0727] Input: Regenerated content file
[0728] Output: Playable content file
[0729] Specific behavior:
[0730] The device receives the regenerated content sent from the server and stores it in local storage. The user can then use MoviePy to play the saved customized content and enjoy watching it.
[0731] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0732] The present invention provides a more interactive and emotional experience by combining a system for personalizing the viewing experience of movies and video content with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0733] User operations
[0734] 1. Launching the application
[0735] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[0736] 2. Face and voice capture
[0737] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0738] When a user faces the camera and speaks into the microphone, the device captures video and audio in real time and records them as digital data.
[0739] 3. Analysis by Emotion Engine
[0740] The captured facial and voice data is input into the emotion engine, which analyzes the user's emotions. The device sends the user's facial expressions and tone of voice to the emotion engine, which then generates emotion data.
[0741] 4. Content Selection
[0742] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[0743] Server Processing
[0744] 1. Data Receipt and Analysis
[0745] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[0746] 2. Acquiring character data
[0747] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0748] 3. Use of generative AI
[0749] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[0750] 4. Generating new lines
[0751] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0752] 5. Building regenerative content
[0753] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[0754] 6. Data Compression and Transfer
[0755] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[0756] Terminal handling
[0757] 1. Receiving new content
[0758] The terminal receives the regenerated content sent from the server and stores it in local storage.
[0759] 2. Playing content
[0760] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0761] Specific examples
[0762] For example, suppose Person B captures his or her face and voice, and the emotion engine detects "joy." If Person B selects an action movie, the server uses Person B's emotional data to adjust the facial expression and tone of voice of the protagonist in the action scene to match joy. Furthermore, the server changes the dialogue of other characters to call Person B's name, generating a new action movie that reflects the emotional expression. This new content is compressed, encoded, and sent to Person B's device. Person B can then play this customized action movie in the application and enjoy a more emotionally resonant viewing experience.
[0763] This system not only allows users to have a more interactive and emotional experience, but also provides security to prevent the generation of inappropriate content.
[0764] The processing flow will be explained below.
[0765] Step 1:
[0766] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[0767] Step 2:
[0768] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0769] Step 3:
[0770] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[0771] Step 4:
[0772] The device sends the captured data to the emotion engine, which analyzes the user's facial expressions and tone of voice to generate emotion data.
[0773] Step 5:
[0774] The device displays a content selection interface, allowing the user to select the movie or video content they want to view.
[0775] Step 6:
[0776] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[0777] Step 7:
[0778] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[0779] Step 8:
[0780] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[0781] Step 9:
[0782] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[0783] Step 10:
[0784] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[0785] Step 11:
[0786] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[0787] Step 12:
[0788] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[0789] Step 13:
[0790] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[0791] Example 2
[0792] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0793] Conventional movie and video content has primarily been a one-way viewing experience, making it difficult to customize it to reflect the user's emotions and individual characteristics. The generation of content that interactively responds to the viewer's emotions has been extremely limited. Furthermore, conventional technology has been inadequate for reflecting the user's face and voice onto characters in movies and videos in real time. By solving these issues, it was necessary to provide a more personal and emotional viewing experience.
[0794] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing the facial expressions and movements of characters in detail from a database, a means for analyzing the line database to obtain line lists of other characters and generate new lines, and a means for replacing the face and voice of a character based on the face and voice data of a user using the generated AI model. This makes it possible to replace the face and voice of a character in selected content with the face and voice of the user, thereby interactively customizing the content according to the user's emotions.
[0795] "User" means an individual who utilizes the System to customize movie and video content.
[0796] "Device" means a device that uses a camera and microphone to capture face and voice, including a smartphone, tablet, or computer.
[0797] "Capture" refers to the process of using a camera and microphone to digitally record facial and audio data.
[0798] "Emotion engine" refers to software that analyzes a user's emotions from captured facial and voice data and generates emotional data.
[0799] "Generative AI model" refers to an artificial intelligence model that replaces the faces and voices of characters in movies and video content with those of the user based on the user's facial and voice data.
[0800] A "dialogue database" refers to a database that stores dialogue used in movies and video content and keeps it in an analyzable format.
[0801] "Secure communication protocol" refers to the communication protocol for securely transmitting captured data to a server.
[0802] "Regenerated Content" refers to customized film or video content generated based on a user's face and voice, emotional data, and new dialogue data.
[0803] "Editing Software" means video editing tools used to create Regenerated Content, including, but not limited to, Adobe Premiere Pro.
[0804] "Compression / encoding" refers to the process of converting digital content to a specific bit rate and format to reduce the data size.
[0805] "Local storage" refers to a memory area that exists within a device and is used to store data.
[0806] The present invention provides a system that provides a more interactive and emotional experience by customizing the viewing experience of movies and video content and combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[0807] Hardware and software used
[0808] This system uses devices such as smartphones, tablets, and PCs. Each device must be equipped with a camera and microphone. The server is a back-end server that receives data using a secure communication protocol (e.g., HTTPS).
[0809] Additionally, the following specific software is used:
[0810] Emotion engine: Emotion recognition algorithm (e.g. Microsoft Azure Emotion API)
[0811] Generative AI models: Generative AI models (e.g., OpenAI GPT-3, DeepFaceLab)
[0812] Editing software: Video editing software (e.g. Adobe Premiere Pro)
[0813] Specific implementation methods
[0814] 1. Launching the application
[0815] The user launches the application by tapping the app icon on the device's home screen, and the application displays its initial screen.
[0816] 2. Face and voice capture
[0817] 1. The device will request permission to access the camera and microphone. If the user allows it, the face and voice capture page will open.
[0818] 2. The user faces the camera and speaks as instructed. The device captures this in real time and generates digital data.
[0819] 3. Analysis by Emotion Engine
[0820] The device sends the captured face and voice data to the emotion engine, which analyzes the data and generates emotion data.
[0821] 4. Content Selection
[0822] The user selects their favorite movie from the "Choose a Movie" interface. Once the user makes a selection, the device displays detailed information from the content library.
[0823] 5. Data processing by the server
[0824] 1. The device sends the capture data and emotion data to the server, which then retrieves the character data for the selected movie from the database.
[0825] 2. The server uses a generative AI model to replace the character based on the user's facial and voice data, adjusting facial expressions and tone of voice based on emotional data.
[0826] 3. The server analyzes the dialogue database and generates new dialogue.
[0827] 4. The server uses editing software such as Adobe Premiere Pro to integrate the new characters, voices, and emotional expressions into the original data.
[0828] 6. Content Compression and Transfer
[0829] The server compresses the regenerated content using an appropriate video codec, such as H.264, and sends it to the device.
[0830] 7. Receiving and Playing Content
[0831] 1. The device receives the regenerated content from the server and stores it in local storage.
[0832] 2. The user selects new content within the application and begins playback.
[0833] Specific examples
[0834] For example, if a user selects an action movie and captures their face and voice, the emotion engine may recognize the user's emotion as "joy." In this case, the server will replace the main character of the movie with the user's own character based on the user's face and voice data, and adjust the characters' facial expressions and tone of voice to reflect "joy." It will also change the dialogue so that other characters call the user's name, creating a customized action movie. This new content is then sent to the user's device, providing the user with an interactive and emotional viewing experience.
[0835] Prompt Sentence Examples
[0836] Example prompts for generative AI models:
[0837] "Based on the user's facial and voice data, replace the face and voice of the main character in an action movie with the user's. Also, if the emotion data is 'joy', adjust the main character's facial expression and tone of voice to be happy, and generate new lines for other characters in the movie to call the user's name."
[0838] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0839] Step 1:
[0840] The user launches an application on the device.
[0841] Specific operation: The user taps the app icon on the device's home screen and the application launches.
[0842] Input: User taps.
[0843] Output: The initial screen of the application is displayed.
[0844] Step 2:
[0845] The device will request permission to access the camera and microphone.
[0846] Specific behavior: The application displays "Please allow access to your camera and microphone." The user taps "Allow."
[0847] Input: User's access permission permission operation.
[0848] Output: The camera and microphone will be activated and the Face and Voice Capture page will open.
[0849] Step 3:
[0850] The user faces the camera and speaks into the microphone, which the device captures as digital data in real time.
[0851] Specific operation: The user follows the instructions to face the camera and say "hello." The device captures the facial image and audio and records them as digital data.
[0852] Input: User's face and voice.
[0853] Output: Digital data of the captured face and voice.
[0854] Step 4:
[0855] The device sends the captured face and voice data to the emotion engine.
[0856] How it works: The device uploads face and voice data to the cloud-based emotion engine, which then analyzes the data and generates emotion data.
[0857] Input: Captured digital face and voice data.
[0858] Output: Parsed emotion data.
[0859] Step 5:
[0860] The user selects the movie or video content they want to view on the device interface.
[0861] What happens: The user selects a movie from the content library and confirms the selection.
[0862] Input: User's movie title selection.
[0863] Output: Detailed information about the selected movie is displayed.
[0864] Step 6:
[0865] The device transmits the capture data and emotion data to the server.
[0866] Specific operation: The device encrypts the data and sends it to the server via a secure communication protocol.
[0867] Input: Capture data, emotion data.
[0868] Output: The server receives face, voice, and emotion data.
[0869] Step 7:
[0870] The server retrieves character data from the database and analyzes it.
[0871] Specific operation: The server accesses a movie database to obtain data on the characters' faces, voices, movements, and lines. The obtained data is then analyzed using a facial recognition algorithm and a voice synthesis model.
[0872] Input: A content database containing character data.
[0873] Output: Analyzed character face, voice, movement, and dialogue data.
[0874] Step 8:
[0875] The server uses a generative AI model to replace characters based on the user's facial and voice data.
[0876] How it works: The server uses a generative AI model to replace the user's face and voice data with the character's face and voice, and also adjusts facial expressions and tone of voice based on emotional data.
[0877] Input: Capture data, emotion data, character data.
[0878] Output: Character data replaced with the user's face and voice.
[0879] Step 9:
[0880] The server analyzes the dialogue database and generates new dialogue.
[0881] How it works: The server accesses the dialogue database and generates new dialogue for other characters to call the user's name. The dialogue generation uses a generative AI model.
[0882] Input: Dialogue database, user's name.
[0883] Output: New dialogue data containing the user's name.
[0884] Step 10:
[0885] The server recreates the original video content.
[0886] Specific operation: The server uses editing software such as Adobe Premiere Pro to integrate the generated face, voice, emotion data, and new dialogue data into the original video and re-edit it.
[0887] Input: Generated face, voice, and emotion data, new dialogue data, and original video content.
[0888] Output: The regenerated video content.
[0889] Step 11:
[0890] The server compresses and encodes the regenerated content and sends it to the device.
[0891] Specific operation: The server compresses the reproduced content using a video codec such as H.264 and sends it to the terminal via a secure communication protocol.
[0892] Input: The regenerated content.
[0893] Output: Compressed and encoded regenerated content.
[0894] Step 12:
[0895] The device receives the regenerated content and stores it in local storage.
[0896] Specific operation: The device receives the regenerated content from the server through a secure communication protocol and stores it in local storage.
[0897] Input: Compressed and encoded regenerated content.
[0898] Output: Regenerated content saved to local storage.
[0899] Step 13:
[0900] The user starts playing the regenerated content within the application on the terminal.
[0901] Specific operation: The user operates the application, selects the regenerated content, and presses the play button. The device's media player starts and the regenerated content is played.
[0902] Input: User playback operations.
[0903] Output: Playback of regenerated content.
[0904] The above processing steps provide the user with a customized and interactive viewing experience.
[0905] (Application example 2)
[0906] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0907] Traditional movie and video content could only provide fixed content and could not be customized according to the viewer's emotions or individual characteristics. This made it difficult to provide a more interactive and emotional viewing experience. Furthermore, it was not possible to reflect the viewer's name or individual characteristics, which resulted in an insufficient personalized experience.
[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into an emotion engine and analyzing it, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for adjusting the character's facial expression and tone of voice based on emotion data, means for changing the lines of other characters to new lines including the user's name, and means for regenerating content based on the replaced data, the adjusted data, and the changed lines and transmitting the regenerated content to the user's terminal. This makes it possible to provide an interactive and emotional viewing experience customized for each user.
[0909] "Means for capturing the user's face and voice in real time" refers to devices or software that instantly record the facial image and voice of the user as they speak to the terminal.
[0910] "Emotion engine" refers to algorithms and software that analyze captured facial and voice data to infer a user's emotional state.
[0911] "Means for replacing the face and voice of a character in content with the user's face and voice based on input data" refers to devices or software that execute a process to replace the face and voice of a character appearing in content using collected data of the user's face and voice.
[0912] "Means for adjusting the facial expressions and tone of voice of characters" refers to devices or software that change the facial expressions and voice characteristics of characters in the content based on emotional data obtained by the emotion engine.
[0913] "Means for changing the lines of other characters into new lines that include the user's name" refers to devices or software that convert the words spoken by characters included in the content into new words that include the user's name.
[0914] "Means for regenerating content" refers to devices or software that regenerate new video and audio data based on the user's face, voice, emotional data, and changed lines.
[0915] "Means for transmitting the regenerated content to the user's device" means the equipment or software that transmits the edited or generated new video and audio to the user's device using the appropriate protocol.
[0916] This invention is a system for customizing a user's viewing experience, implemented using a smartphone application. Specifically, it captures the user's face and voice, analyzes the data with an emotion engine, and uses a generative AI model to customize the faces, voices, emotions, and lines of characters in video content in real time, providing an interactive and emotional viewing experience.
[0917] Hardware and Software Use
[0918] Device: A smartphone is used. The smartphone's front camera and microphone are used to capture face and voice. The smartphone must also have a media player installed for viewing.
[0919] Server: Runs the emotion engine, generative AI model, and dialogue conversion algorithm.
[0920] Emotion engine: Analyzes captured facial and voice data to interpret user emotions in real time. Specific software used includes facial recognition algorithms (e.g., OpenCV) and speech recognition software (e.g., the speech_recognition library).
[0921] Generative AI models: Generate the faces, voices, and expressions of characters in content based on emotional and captured data. Specific examples include face-swapping algorithms and voice synthesis models (e.g., DeepFake, Tacotron).
[0922] Dialogue conversion algorithm: An algorithm to change the dialogue of other characters into new dialogue that includes the user's name. New dialogue is generated using a natural language processing model (e.g., GPT-3).
[0923] System flow
[0924] 1. User operations
[0925] A user launches a smartphone application and uses the camera and microphone to capture their face and voice.
[0926] 2. Data Analysis
[0927] The captured face and voice data is sent to the emotion engine, which analyzes the user's emotional state and formats the results as emotion data, including the user's facial expressions and tone of voice.
[0928] 3. Server Processing
[0929] The server receives the emotion data and inputs it into a generative AI model. The face and voice of the characters in the content are replaced with the user's, and the facial expressions and tone of voice are adjusted. The lines of other characters are also converted into new lines that include the user's name.
[0930] 4. Regenerating and Submitting Content
[0931] The server regenerates the video content based on the adjusted data and the new dialogue. This new content is compressed, encoded, and sent to the user's smartphone using a secure communication protocol.
[0932] 5. Playing Content
[0933] The user's smartphone receives the regenerated content and launches a media player to play it.
[0934] Adding specific examples
[0935] For example, a user opens the app and captures their face and voice. The emotion engine then detects "happiness." Based on this result, the facial expression and voice of the protagonist of the action movie selected by the user are adjusted to reflect a state of happiness. Furthermore, the dialogue of other characters is changed to call the user's name, "Takashi."
[0936] Example prompts for generative AI models
[0937] text
[0938] "User emotion analysis data": "Joy",
[0939] "Character Layer": {
[0940] "Expression": "Joy",
[0941] "Tone": "bright",
[0942] "Name": "Takashi"
[0943] },
[0944] "Dialogue Database": "Generate new dialogue"
[0945] The customized content thus generated is then sent to the user's smartphone, allowing the user to enjoy a more emotionally resonant viewing experience.
[0946] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0947] Step 1:
[0948] A user launches a smartphone application. The user taps the app icon on the device's home screen to open it. The application requests permission to access the camera and microphone and displays a face and voice capture page. The input is the user's operation, and the output is the activation of the camera and microphone.
[0949] Step 2:
[0950] The user faces the camera and speaks into the microphone. The device captures video and records audio in real time. The captured data is temporarily stored in the device's local storage. The input is the user's face and voice, and the output is facial image data and audio data.
[0951] Step 3:
[0952] The device sends the captured facial image data and voice data to the emotion engine to analyze the emotional state. The emotion engine uses facial recognition algorithms and voice analysis algorithms (e.g., OpenCV, speech_recognition) to analyze the user's facial expressions and tone of voice. The input is facial image data and voice data, and the output is the user's emotional data (e.g., "happiness," "excitement," etc.).
[0953] Step 4:
[0954] The device generates emotion data and transmits it to the server along with the captured data using a secure communication protocol (e.g., HTTPS). The input is facial image data, voice data, and emotion data, and the output is the securely transmitted data.
[0955] Step 5:
[0956] The server inputs the received facial image data, voice data, and emotion data into a generative AI model. The generative AI model then replaces the faces and voices of characters in the content with the user's face and voice. The input is the captured data and emotion data, and the output is the converted face and voice of the characters.
[0957] Step 6:
[0958] The server adjusts the facial expressions and tone of voice of the characters in the content based on the emotional data. The generative AI model changes the facial expressions and voice of the characters in real time to match the user's emotions. The input is emotional data, and the output is adjusted facial expression data and tone of voice data.
[0959] Step 7:
[0960] The server analyzes the dialogue of other characters and converts it into new dialogue that includes the user's name. It generates the dialogue using a natural language processing algorithm (e.g., GPT-3). The input is the existing dialogue data, and the output is the new dialogue data.
[0961] Step 8:
[0962] The server regenerates the content based on the facial data, voice data, adjusted facial expression data, and new dialogue data. This data is then integrated using video editing software (e.g., Adobe Premiere Pro). The input is all the converted data, and the output is the regenerated content.
[0963] Step 9:
[0964] The server compresses and encodes the regenerated content, encoding it at the appropriate bitrate, and sends the compressed data to the device using a secure communication protocol. The input is the regenerated content, and the output is the compressed and encoded content.
[0965] Step 10:
[0966] The device receives the regenerated content and saves it to local storage. It then launches the device's media player and begins playing the new content. The input is the compressed and encoded content, and the output is a customized video played to the user.
[0967] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0968] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0969] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0970] [Third embodiment]
[0971] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0972] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0973] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0974] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0975] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0976] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0977] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0978] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0979] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0980] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0981] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0982] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0983] The present invention is a system for personalizing the viewing experience of movies and video content. Users can use their own face and voice to replace characters in the content with themselves. This system is characterized by exchanging data using a secure communication protocol and playing the generated customized content on the user's terminal. Specific embodiments of the system are described below.
[0984] User operations
[0985] 1. Launching the application
[0986] The user launches an application on the device by tapping the app icon on the device's home screen.
[0987] 2. Face and voice capture
[0988] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[0989] When a user faces the camera and speaks into the microphone for a few seconds, the device captures and records video and audio in real time.
[0990] 3. Content Selection
[0991] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[0992] Server Processing
[0993] 1. Data Receipt and Analysis
[0994] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[0995] 2. Acquiring character data
[0996] The server retrieves character data (face, voice, movements, lines) from the database of the selected content.
[0997] 3. Use of generative AI
[0998] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the inputted face and voice data of the user.
[0999] 4. Generating new lines
[1000] The server changes the lines that other characters use to call out to the user to new lines that include the user's name, so that other characters will call out the user's name.
[1001] 5. Building regenerative content
[1002] The server reconstructs the original video content based on the generated facial, voice, and dialogue data, and then uses editing software to integrate the new characters and voices into the original data.
[1003] 6. Data Compression and Transfer
[1004] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[1005] Terminal handling
[1006] 1. Receiving new content
[1007] The terminal receives the regenerated content sent from the server and stores it in local storage.
[1008] 2. Playing content
[1009] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1010] Specific examples
[1011] For example, suppose Person A captures his or her face and voice and selects an animated movie. The server inputs Person A's face and voice into the generation AI, which replaces the face and voice of the main character in the animated movie with Person A's. It also changes the lines of other characters to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can play the content in an application and enjoy the animated movie starring himself or herself as the main character.
[1012] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent the generation of inappropriate content.
[1013] The processing flow will be explained below.
[1014] Step 1:
[1015] The user launches an application on the device by tapping the corresponding app icon on the device's home screen.
[1016] Step 2:
[1017] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1018] Step 3:
[1019] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[1020] Step 4:
[1021] The user selects the movie or video content they want to watch through the device's interface, and the device selects the title from its content library.
[1022] Step 5:
[1023] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a face recognition algorithm and a voice synthesis model.
[1024] Step 6:
[1025] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1026] Step 7:
[1027] The server uses generative AI to replace the face and voice of the characters in the content with those of the user based on the inputted face and voice data of the user, and generates a digital model that reflects the user's characteristics.
[1028] Step 8:
[1029] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1030] Step 9:
[1031] The server recreates the original video content based on the generated facial, voice, and dialogue data, and uses editing software to integrate the new characters and voices into the original data.
[1032] Step 10:
[1033] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[1034] Step 11:
[1035] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[1036] Step 12:
[1037] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1038] Example 1
[1039] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1040] In content such as videos and movies, it is desirable for users to replace characters with themselves to have a more personal and immersive experience. However, this process is technically complex, and there has been no system that makes it easy for users to do so. In addition, security measures to prevent inappropriate content generation are often lacking, so a means of safely providing customized content has been sought.
[1041] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1042] In this invention, the server includes means for transmitting the captured face and voice data to the server using a secure communication protocol, means for replacing the faces and voices of characters in the content with the face and voice of the user based on the input data, and means for changing the lines of other characters to new lines including the user's name, thereby enabling the user to safely and easily replace themselves with characters in the content and enjoy a personal and immersive experience.
[1043] "Means for capturing a user's face and voice in real time" refers to a system that uses a camera and a microphone to record and record a user's facial expressions and voice in real time.
[1044] "Means for inputting captured facial and audio data into generative AI" refers to the process or interface for inputting facial images and audio data captured in real time into an AI model.
[1045] "Means of replacing the faces and voices of characters in content with the face and voice of the user" refers to technology that converts and replaces the faces and voices of specific characters in content based on acquired face and voice data of the user.
[1046] The "means for changing the lines of other characters to new lines that include the user's name" is a text processing algorithm for modifying the lines spoken by other characters to include the user's name.
[1047] "Means for regenerating content" refers to methods or tools for editing and recomposing the original content based on the replaced face or voice data.
[1048] "Means for compressing and encoding the reproduced content" means techniques for reducing the size of the reproduced video and audio data and converting it into an appropriate format.
[1049] The "means for transmitting to the user's terminal" refers to a communication means for safely and quickly transmitting data generated from the server side to the user's device.
[1050] The "means for the user's terminal to play the regenerated content" refers to software or playback functions that allow the user's device to properly play the acquired customized content.
[1051] The present invention is a system that allows users to replace themselves with characters in movies and video content to enjoy a personalized viewing experience. The system includes a means for capturing the user's face and voice in real time and using a generative AI model to replace the face and voice of characters in the content with that of the user.
[1052] First, the user launches the application installed on their device, which can be a smartphone or tablet. The user grants camera and microphone access permissions from the application's main screen and captures their face and voice. This captured data is encoded in real time and sent to the server using a secure communication protocol (e.g., HTTPS).
[1053] The server uses a facial recognition algorithm (e.g., OpenCV) and a voice synthesis model (e.g., WaveNet) to analyze the received face and voice data. The server simultaneously retrieves the face, voice, movement, and dialogue data of the character from a database of the selected content. A generative AI model (e.g., DeepFake or DALL-E) is then used to replace the character's face and voice with the user's.
[1054] Additionally, a text processing algorithm is run to replace dialogue spoken by other characters with new dialogue that includes the user's name, causing other characters to call out the user's name. The server then uses this processed data to reconstruct the original video content in editing software such as Adobe Premiere Pro or Final Cut Pro. During this editing process, the new faces, voices, and dialogue are integrated into the timeline to create a seamless visual experience.
[1055] The regenerated content is compressed and encoded (e.g., H.264 codec) to optimize file size, and then transmitted to the device using a secure communication protocol. The user can then play the regenerated content stored in the device's local storage using the application's media player and enjoy customized movies and videos featuring themselves.
[1056] As a concrete example, suppose Person A captures his or her face and voice using an app and selects his or her favorite animated movie. The system inputs Person A's face and voice into a generative AI model, which then replaces the face and voice of the main character in the animated movie with Person A's. The dialogue of other characters in the anime is also changed to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie in which he or she appears as the main character.
[1057] Example prompt sentence:
[1058] "Mr. A, please face the camera to capture your face and speak into the microphone. Then, choose your favorite movie and press the start button."
[1059] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent inappropriate content from being generated.
[1060] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1061] Step 1:
[1062] Input: A user launches an application by tapping the app icon on the device's home screen.
[1063] Processing: The user launches the application by tapping the app icon on the device. The app displays a dialog requesting permission to use the camera and microphone.
[1064] Output: The application launches and displays a screen requesting camera and microphone access permissions.
[1065] What happens: The user taps the button to allow access permissions, granting access to the camera and microphone.
[1066] Step 2:
[1067] Input: The user faces the camera and speaks into the microphone for a few seconds.
[1068] Processing: The device captures the user's face and voice in real time and encodes it.
[1069] Output: The captured face and audio data is temporarily stored on the device.
[1070] What happens: The user looks at the camera and speaks into the microphone as instructed, and the application records and films this.
[1071] Step 3:
[1072] Input: Captured face and audio data.
[1073] Processing: The device sends the captured data to the server using a secure communication protocol (e.g., HTTPS).
[1074] Output: Face and voice data received by the server.
[1075] What it does: The application encrypts the captured data and sends it over the internet to a server.
[1076] Step 4:
[1077] Input: Face and voice data received by the server.
[1078] Processing: The server analyzes the data using facial recognition algorithms (e.g., OpenCV) and speech synthesis models (e.g., WaveNet).
[1079] Output: Face and voice data analyzed on the server.
[1080] Specific operation: The server analyzes facial features and voice waveforms to extract the necessary data.
[1081] Step 5:
[1082] Input: Analyzed face and audio data, selected content.
[1083] Processing: The server retrieves the character's face, voice, movement, and dialogue data from the database of the selected content.
[1084] Output: Character face data, voice data, movement data, and dialogue data.
[1085] What happens: The server queries the database for the specified content and retrieves the required data.
[1086] Step 6:
[1087] Input: Captured character data and parsed user data.
[1088] Processing: The server uses a generative AI model (e.g., DeepFake or DALL-E) to replace the character's face and voice with the user's face and voice.
[1089] Output: The replaced data.
[1090] Specific operation: The server runs the AI model and synthesizes the characters' faces and voices with those of the user.
[1091] Step 7:
[1092] Input: Replaced face and voice data and original description data.
[1093] What happens: The server changes the dialogue of the other characters to new dialogue that includes the user's name.
[1094] Output: The corrected dialogue data.
[1095] What it does: It uses a text processing algorithm to change the dialogue of characters to match the user's name.
[1096] Step 8:
[1097] Input: Replaced face and voice data and modified dialogue data.
[1098] Processing: The server uses editing software such as Adobe Premiere Pro or Final Cut Pro to reconstruct the content.
[1099] Output: The regenerated content.
[1100] How it works: The server synthesizes new data on the editing software's timeline to create a seamless video.
[1101] Step 9:
[1102] Input: The regenerated content.
[1103] Processing: The server compresses the regenerated content and encodes it with the H.264 codec.
[1104] Output: Compressed and encoded content data.
[1105] Specific operation: The server uses encoding software to optimize the file size and sends it to the device using a secure communication protocol.
[1106] Step 10:
[1107] Input: Compressed and encoded content data sent to the terminal.
[1108] Processing: The device saves the received content in local storage.
[1109] Output: Content saved to local storage.
[1110] Specific operation: The device receives the data and automatically saves it in the specified folder.
[1111] Step 11:
[1112] Input: Content stored in local storage.
[1113] Action: The device launches a media player and plays the stored content.
[1114] Output: The customized content that is played.
[1115] What happens: The user taps the play button in the app and starts watching the customized movie or video.
[1116] (Application example 1)
[1117] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1118] Traditional movie and video content viewing experiences lack personalization and immersion for viewers. It is difficult for viewers to customize the experience by replacing characters with their own faces and voices, providing a limited experience. Furthermore, there are insufficient means to securely and efficiently transmit this customized content.
[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1120] In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into a generative AI model, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for regenerating the content based on the replaced data, means for transmitting the regenerated content to the user's terminal, means for the user to select content they wish to view, and means for saving the content in local storage and playing it back. This allows the user to securely and efficiently receive customized content in which they themselves are replaced with characters, and enjoy a highly immersive viewing experience.
[1121] "User" refers to the person who interacts with the system, captures their face and voice, and views customized content.
[1122] "Real-time face and voice capture method" refers to a device or software method for capturing a user's face and voice and collecting that data.
[1123] "Generative AI model" refers to an artificial intelligence model that uses captured facial and voice data to generate the face and voice of a new character.
[1124] "Input means" refers to the mechanism or method for inputting captured facial and voice data into a generative AI model.
[1125] "Characters" refers to characters that appear in movies and video content.
[1126] "Means of replacement" refers to the operation of using a generative AI model to replace the faces and voices of characters in the content with the user's face and voice.
[1127] "Regeneration means" refers to the method or process for creating new edited content based on replaced data.
[1128] "Means for playing content" refers to the mechanism or method for playing the created customized content on the user's terminal.
[1129] "Secure communication protocol" refers to a communication method for securely transmitting captured data to a server.
[1130] "Compression / Encoding Methods" means the methods and techniques that convert the data in the Reproduced Content into a format that can be efficiently transferred.
[1131] "Local storage" refers to the data storage area built into the user's device.
[1132] MODE FOR CARRYING OUT THE INVENTION
[1133] System program and processing overview
[1134] An embodiment of the present invention is a system that captures a user's face and voice to provide customized content. This system operates in cooperation with a terminal, such as a smartphone, smart glasses, a head-mounted display, or a robot, and a server. The specific operation of the system and its implementation method are described below.
[1135] Hardware and software configuration:
[1136] 1. Device:
[1137] Camera: Captures the user's face using OpenCV.
[1138] Microphone: Uses a SoundDevice to capture the user's voice.
[1139] Storage: Save data to local storage.
[1140] Communication: Use Requests to transmit data securely.
[1141] Media Player: Use MoviePy to play customized content.
[1142] 2. Server:
[1143] Generative AI model: Replaces characters in content based on the user's facial and voice data.
[1144] Database: Stores data on movies and video content.
[1145] Communication protocol: Use a communication protocol to securely transmit data.
[1146] Encoding software: compresses and encodes the reproduced content and sends it to the user's device.
[1147] Data processing and calculation flow:
[1148] 1. On the user's device:
[1149] The user launches an application on the device to capture their face and voice. The captured face image and voice data are temporarily stored in local storage.
[1150] The user selects the movie or video content they want to watch, and the content selection information is also stored in local storage.
[1151] 2. Server:
[1152] The facial and voice data sent from the device is received and input into the generative AI model, and this process uses a secure communication protocol to protect the data.
[1153] The generative AI model replaces the faces and voices of characters in the content with those of the user based on the input data.
[1154] Once all processing is complete, the regenerated content is compressed and encoded using encoding software and sent to the user's device.
[1155] 3. Playing content:
[1156] The user's device receives the new content and saves it to local storage, where the user can play the saved customized content using MoviePy.
[1157] Examples:
[1158] For example, if a user captures their face and voice and selects a particular movie, they can send the data to a server using a prompt like this:
[1159] bash
[1160] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[1161] This prompt example uploads the files face.jpg and audio.wav to the server and specifies the selected movie ID (example_movie_id). The server generates customized content based on the received data and sends it to the user's device.
[1162] As described above, by using this system, users can have a personalized experience by replacing themselves with characters in movies and video content based on their facial and voice data.
[1163] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1164] Step 1:
[1165] The user launches an application on the device. The user opens the application by tapping the app icon on the device's home screen. The user then proceeds to the face and voice capture screen.
[1166] Input: Face and voice capture request
[1167] Output: Face image (face.jpg) and audio data (audio.wav)
[1168] Specific behavior:
[1169] The device will request permission to access the camera and microphone and display a face and audio capture page. When the user faces the camera and speaks into the microphone, the device will capture video and audio in real time and save them to local storage as "face.jpg" and "audio.wav."
[1170] Step 2:
[1171] The user selects the movie or video content they want to watch by using the application's interface to select the title from the content library.
[1172] Input: User-selected content ID
[1173] Output: Selected content ID (content_id)
[1174] Specific behavior:
[1175] The device interface displays a content selection screen to the user, and the user selects the content they want to watch by tapping on it. This selection information is saved as "content_id."
[1176] Step 3:
[1177] The device transmits the captured face and voice data and the selected content ID to a server using a secure communication protocol.
[1178] Input: Face image (face.jpg), audio data (audio.wav), content ID (content_id)
[1179] Output: Send data to the server
[1180] Specific behavior:
[1181] The terminal uses the Requests library to send data to the server using a secure communication protocol, specifically using the following prompt sentence:
[1182] bash
[1183] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[1184] Step 4:
[1185] The server analyzes the received facial and voice data and inputs it into a generative AI model, which then replaces the face and voice of the characters in the content with those of the user based on their face and voice.
[1186] Input: Face image (face.jpg), audio data (audio.wav)
[1187] Output: Replaced face and voice data
[1188] Specific behavior:
[1189] The server uses a generative AI model to analyze the user's facial image and voice data and process it to replace them with a character in the content, converting the character's face and voice to that of the user.
[1190] Step 5:
[1191] The server regenerates the content based on the replaced face and voice data, compresses and encodes the regenerated content, and sends it to the user's device.
[1192] Input: Replaced face and voice data
[1193] Output: Regenerated content files
[1194] Specific behavior:
[1195] The server uses encoding software to compress and encode the reproduced content into an efficiently transferable format, and the encoded content file is sent to the user's device using a secure communications protocol.
[1196] Step 6:
[1197] The device receives the regenerated content and stores it in local storage, allowing the user to play and enjoy the content.
[1198] Input: Regenerated content file
[1199] Output: Playable content file
[1200] Specific behavior:
[1201] The device receives the regenerated content sent from the server and stores it in local storage. The user can then use MoviePy to play the saved customized content and enjoy watching it.
[1202] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1203] The present invention provides a more interactive and emotional experience by combining a system for personalizing the viewing experience of movies and video content with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[1204] User operations
[1205] 1. Launching the application
[1206] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[1207] 2. Face and voice capture
[1208] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1209] When a user faces the camera and speaks into the microphone, the device captures video and audio in real time and records them as digital data.
[1210] 3. Analysis by Emotion Engine
[1211] The captured facial and voice data is input into the emotion engine, which analyzes the user's emotions. The device sends the user's facial expressions and tone of voice to the emotion engine, which then generates emotion data.
[1212] 4. Content Selection
[1213] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[1214] Server Processing
[1215] 1. Data Receipt and Analysis
[1216] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[1217] 2. Acquiring character data
[1218] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1219] 3. Use of generative AI
[1220] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[1221] 4. Generating new lines
[1222] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1223] 5. Building regenerative content
[1224] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[1225] 6. Data Compression and Transfer
[1226] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[1227] Terminal handling
[1228] 1. Receiving new content
[1229] The terminal receives the regenerated content sent from the server and stores it in local storage.
[1230] 2. Playing content
[1231] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1232] Specific examples
[1233] For example, suppose Person B captures his or her face and voice, and the emotion engine detects "joy." If Person B selects an action movie, the server uses Person B's emotional data to adjust the facial expression and tone of voice of the protagonist in the action scene to match joy. Furthermore, the server changes the dialogue of other characters to call Person B's name, generating a new action movie that reflects the emotional expression. This new content is compressed, encoded, and sent to Person B's device. Person B can then play this customized action movie in the application and enjoy a more emotionally resonant viewing experience.
[1234] This system not only allows users to have a more interactive and emotional experience, but also provides security to prevent the generation of inappropriate content.
[1235] The processing flow will be explained below.
[1236] Step 1:
[1237] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[1238] Step 2:
[1239] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1240] Step 3:
[1241] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[1242] Step 4:
[1243] The device sends the captured data to the emotion engine, which analyzes the user's facial expressions and tone of voice to generate emotion data.
[1244] Step 5:
[1245] The device displays a content selection interface, allowing the user to select the movie or video content they want to view.
[1246] Step 6:
[1247] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[1248] Step 7:
[1249] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1250] Step 8:
[1251] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[1252] Step 9:
[1253] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1254] Step 10:
[1255] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[1256] Step 11:
[1257] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[1258] Step 12:
[1259] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[1260] Step 13:
[1261] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1262] Example 2
[1263] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1264] Conventional movie and video content has primarily been a one-way viewing experience, making it difficult to customize it to reflect the user's emotions and individual characteristics. The generation of content that interactively responds to the viewer's emotions has been extremely limited. Furthermore, conventional technology has been inadequate for reflecting the user's face and voice onto characters in movies and videos in real time. By solving these issues, it was necessary to provide a more personal and emotional viewing experience.
[1265] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing the facial expressions and movements of characters in detail from a database, a means for analyzing the line database to obtain line lists of other characters and generate new lines, and a means for replacing the face and voice of a character based on the face and voice data of a user using the generated AI model. This makes it possible to replace the face and voice of a character in selected content with the face and voice of the user, thereby interactively customizing the content according to the user's emotions.
[1266] "User" means an individual who utilizes the System to customize movie and video content.
[1267] "Device" means a device that uses a camera and microphone to capture face and voice, including a smartphone, tablet, or computer.
[1268] "Capture" refers to the process of using a camera and microphone to digitally record facial and audio data.
[1269] "Emotion engine" refers to software that analyzes a user's emotions from captured facial and voice data and generates emotional data.
[1270] "Generative AI model" refers to an artificial intelligence model that replaces the faces and voices of characters in movies and video content with those of the user based on the user's facial and voice data.
[1271] A "dialogue database" refers to a database that stores dialogue used in movies and video content and keeps it in an analyzable format.
[1272] "Secure communication protocol" refers to the communication protocol for securely transmitting captured data to a server.
[1273] "Regenerated Content" refers to customized film or video content generated based on a user's face and voice, emotional data, and new dialogue data.
[1274] "Editing Software" means video editing tools used to create Regenerated Content, including, but not limited to, Adobe Premiere Pro.
[1275] "Compression / encoding" refers to the process of converting digital content to a specific bit rate and format to reduce the data size.
[1276] "Local storage" refers to a memory area that exists within a device and is used to store data.
[1277] The present invention provides a system that provides a more interactive and emotional experience by customizing the viewing experience of movies and video content and combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[1278] Hardware and software used
[1279] This system uses devices such as smartphones, tablets, and PCs. Each device must be equipped with a camera and microphone. The server is a back-end server that receives data using a secure communication protocol (e.g., HTTPS).
[1280] Additionally, the following specific software is used:
[1281] Emotion engine: Emotion recognition algorithm (e.g. Microsoft Azure Emotion API)
[1282] Generative AI models: Generative AI models (e.g., OpenAI GPT-3, DeepFaceLab)
[1283] Editing software: Video editing software (e.g. Adobe Premiere Pro)
[1284] Specific implementation methods
[1285] 1. Launching the application
[1286] The user launches the application by tapping the app icon on the device's home screen, and the application displays its initial screen.
[1287] 2. Face and voice capture
[1288] 1. The device will request permission to access the camera and microphone. If the user allows it, the face and voice capture page will open.
[1289] 2. The user faces the camera and speaks as instructed. The device captures this in real time and generates digital data.
[1290] 3. Analysis by Emotion Engine
[1291] The device sends the captured face and voice data to the emotion engine, which analyzes the data and generates emotion data.
[1292] 4. Content Selection
[1293] The user selects their favorite movie from the "Choose a Movie" interface. Once the user makes a selection, the device displays detailed information from the content library.
[1294] 5. Data processing by the server
[1295] 1. The device sends the capture data and emotion data to the server, which then retrieves the character data for the selected movie from the database.
[1296] 2. The server uses a generative AI model to replace the character based on the user's facial and voice data, adjusting facial expressions and tone of voice based on emotional data.
[1297] 3. The server analyzes the dialogue database and generates new dialogue.
[1298] 4. The server uses editing software such as Adobe Premiere Pro to integrate the new characters, voices, and emotional expressions into the original data.
[1299] 6. Content Compression and Transfer
[1300] The server compresses the regenerated content using an appropriate video codec, such as H.264, and sends it to the device.
[1301] 7. Receiving and Playing Content
[1302] 1. The device receives the regenerated content from the server and stores it in local storage.
[1303] 2. The user selects new content within the application and begins playback.
[1304] Specific examples
[1305] For example, if a user selects an action movie and captures their face and voice, the emotion engine may recognize the user's emotion as "joy." In this case, the server will replace the main character of the movie with the user's own character based on the user's face and voice data, and adjust the characters' facial expressions and tone of voice to reflect "joy." It will also change the dialogue so that other characters call the user's name, creating a customized action movie. This new content is then sent to the user's device, providing the user with an interactive and emotional viewing experience.
[1306] Prompt Sentence Examples
[1307] Example prompts for generative AI models:
[1308] "Based on the user's facial and voice data, replace the face and voice of the main character in an action movie with the user's. Also, if the emotion data is 'joy', adjust the main character's facial expression and tone of voice to be happy, and generate new lines for other characters in the movie to call the user's name."
[1309] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1310] Step 1:
[1311] The user launches an application on the device.
[1312] Specific operation: The user taps the app icon on the device's home screen and the application launches.
[1313] Input: User taps.
[1314] Output: The initial screen of the application is displayed.
[1315] Step 2:
[1316] The device will request permission to access the camera and microphone.
[1317] Specific behavior: The application displays "Please allow access to your camera and microphone." The user taps "Allow."
[1318] Input: User's access permission permission operation.
[1319] Output: The camera and microphone will be activated and the Face and Voice Capture page will open.
[1320] Step 3:
[1321] The user faces the camera and speaks into the microphone, which the device captures as digital data in real time.
[1322] Specific operation: The user follows the instructions to face the camera and say "hello." The device captures the facial image and audio and records them as digital data.
[1323] Input: User's face and voice.
[1324] Output: Digital data of the captured face and voice.
[1325] Step 4:
[1326] The device sends the captured face and voice data to the emotion engine.
[1327] How it works: The device uploads face and voice data to the cloud-based emotion engine, which then analyzes the data and generates emotion data.
[1328] Input: Captured digital face and voice data.
[1329] Output: Parsed emotion data.
[1330] Step 5:
[1331] The user selects the movie or video content they want to view on the device interface.
[1332] What happens: The user selects a movie from the content library and confirms the selection.
[1333] Input: User's movie title selection.
[1334] Output: Detailed information about the selected movie is displayed.
[1335] Step 6:
[1336] The device transmits the capture data and emotion data to the server.
[1337] Specific operation: The device encrypts the data and sends it to the server via a secure communication protocol.
[1338] Input: Capture data, emotion data.
[1339] Output: The server receives face, voice, and emotion data.
[1340] Step 7:
[1341] The server retrieves character data from the database and analyzes it.
[1342] Specific operation: The server accesses a movie database to obtain data on the characters' faces, voices, movements, and lines. The obtained data is then analyzed using a facial recognition algorithm and a voice synthesis model.
[1343] Input: A content database containing character data.
[1344] Output: Analyzed character face, voice, movement, and dialogue data.
[1345] Step 8:
[1346] The server uses a generative AI model to replace characters based on the user's facial and voice data.
[1347] How it works: The server uses a generative AI model to replace the user's face and voice data with the character's face and voice, and also adjusts facial expressions and tone of voice based on emotional data.
[1348] Input: Capture data, emotion data, character data.
[1349] Output: Character data replaced with the user's face and voice.
[1350] Step 9:
[1351] The server analyzes the dialogue database and generates new dialogue.
[1352] How it works: The server accesses the dialogue database and generates new dialogue for other characters to call the user's name. The dialogue generation uses a generative AI model.
[1353] Input: Dialogue database, user's name.
[1354] Output: New dialogue data containing the user's name.
[1355] Step 10:
[1356] The server recreates the original video content.
[1357] Specific operation: The server uses editing software such as Adobe Premiere Pro to integrate the generated face, voice, emotion data, and new dialogue data into the original video and re-edit it.
[1358] Input: Generated face, voice, and emotion data, new dialogue data, and original video content.
[1359] Output: The regenerated video content.
[1360] Step 11:
[1361] The server compresses and encodes the regenerated content and sends it to the device.
[1362] Specific operation: The server compresses the reproduced content using a video codec such as H.264 and sends it to the terminal via a secure communication protocol.
[1363] Input: The regenerated content.
[1364] Output: Compressed and encoded regenerated content.
[1365] Step 12:
[1366] The device receives the regenerated content and stores it in local storage.
[1367] Specific operation: The device receives the regenerated content from the server through a secure communication protocol and stores it in local storage.
[1368] Input: Compressed and encoded regenerated content.
[1369] Output: Regenerated content saved to local storage.
[1370] Step 13:
[1371] The user starts playing the regenerated content within the application on the terminal.
[1372] Specific operation: The user operates the application, selects the regenerated content, and presses the play button. The device's media player starts and the regenerated content is played.
[1373] Input: User playback operations.
[1374] Output: Playback of regenerated content.
[1375] The above processing steps provide the user with a customized and interactive viewing experience.
[1376] (Application example 2)
[1377] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1378] Traditional movie and video content could only provide fixed content and could not be customized according to the viewer's emotions or individual characteristics. This made it difficult to provide a more interactive and emotional viewing experience. Furthermore, it was not possible to reflect the viewer's name or individual characteristics, which resulted in an insufficient personalized experience.
[1379] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into an emotion engine and analyzing it, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for adjusting the character's facial expression and tone of voice based on emotion data, means for changing the lines of other characters to new lines including the user's name, and means for regenerating content based on the replaced data, the adjusted data, and the changed lines and transmitting the regenerated content to the user's terminal. This makes it possible to provide an interactive and emotional viewing experience customized for each user.
[1380] "Means for capturing the user's face and voice in real time" refers to devices or software that instantly record the facial image and voice of the user as they speak to the terminal.
[1381] "Emotion engine" refers to algorithms and software that analyze captured facial and voice data to infer a user's emotional state.
[1382] "Means for replacing the face and voice of a character in content with the user's face and voice based on input data" refers to devices or software that execute a process to replace the face and voice of a character appearing in content using collected data of the user's face and voice.
[1383] "Means for adjusting the facial expressions and tone of voice of characters" refers to devices or software that change the facial expressions and voice characteristics of characters in the content based on emotional data obtained by the emotion engine.
[1384] "Means for changing the lines of other characters into new lines that include the user's name" refers to devices or software that convert the words spoken by characters included in the content into new words that include the user's name.
[1385] "Means for regenerating content" refers to devices or software that regenerate new video and audio data based on the user's face, voice, emotional data, and changed lines.
[1386] "Means for transmitting the regenerated content to the user's device" means the equipment or software that transmits the edited or generated new video and audio to the user's device using the appropriate protocol.
[1387] This invention is a system for customizing a user's viewing experience, implemented using a smartphone application. Specifically, it captures the user's face and voice, analyzes the data with an emotion engine, and uses a generative AI model to customize the faces, voices, emotions, and lines of characters in video content in real time, providing an interactive and emotional viewing experience.
[1388] Hardware and Software Use
[1389] Device: A smartphone is used. The smartphone's front camera and microphone are used to capture face and voice. The smartphone must also have a media player installed for viewing.
[1390] Server: Runs the emotion engine, generative AI model, and dialogue conversion algorithm.
[1391] Emotion engine: Analyzes captured facial and voice data to interpret user emotions in real time. Specific software used includes facial recognition algorithms (e.g., OpenCV) and speech recognition software (e.g., the speech_recognition library).
[1392] Generative AI models: Generate the faces, voices, and expressions of characters in content based on emotional and captured data. Specific examples include face-swapping algorithms and voice synthesis models (e.g., DeepFake, Tacotron).
[1393] Dialogue conversion algorithm: An algorithm to change the dialogue of other characters into new dialogue that includes the user's name. New dialogue is generated using a natural language processing model (e.g., GPT-3).
[1394] System flow
[1395] 1. User operations
[1396] A user launches a smartphone application and uses the camera and microphone to capture their face and voice.
[1397] 2. Data Analysis
[1398] The captured face and voice data is sent to the emotion engine, which analyzes the user's emotional state and formats the results as emotion data, including the user's facial expressions and tone of voice.
[1399] 3. Server Processing
[1400] The server receives the emotion data and inputs it into a generative AI model. The face and voice of the characters in the content are replaced with the user's, and the facial expressions and tone of voice are adjusted. The lines of other characters are also converted into new lines that include the user's name.
[1401] 4. Regenerating and Submitting Content
[1402] The server regenerates the video content based on the adjusted data and the new dialogue. This new content is compressed, encoded, and sent to the user's smartphone using a secure communication protocol.
[1403] 5. Playing Content
[1404] The user's smartphone receives the regenerated content and launches a media player to play it.
[1405] Adding specific examples
[1406] For example, a user opens the app and captures their face and voice. The emotion engine then detects "happiness." Based on this result, the facial expression and voice of the protagonist of the action movie selected by the user are adjusted to reflect a state of happiness. Furthermore, the dialogue of other characters is changed to call the user's name, "Takashi."
[1407] Example prompts for generative AI models
[1408] text
[1409] "User emotion analysis data": "Joy",
[1410] "Character Layer": {
[1411] "Expression": "Joy",
[1412] "Tone": "bright",
[1413] "Name": "Takashi"
[1414] },
[1415] "Dialogue Database": "Generate new dialogue"
[1416] The customized content thus generated is then sent to the user's smartphone, allowing the user to enjoy a more emotionally resonant viewing experience.
[1417] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1418] Step 1:
[1419] A user launches a smartphone application. The user taps the app icon on the device's home screen to open it. The application requests permission to access the camera and microphone and displays a face and voice capture page. The input is the user's operation, and the output is the activation of the camera and microphone.
[1420] Step 2:
[1421] The user faces the camera and speaks into the microphone. The device captures video and records audio in real time. The captured data is temporarily stored in the device's local storage. The input is the user's face and voice, and the output is facial image data and audio data.
[1422] Step 3:
[1423] The device sends the captured facial image data and voice data to the emotion engine to analyze the emotional state. The emotion engine uses facial recognition algorithms and voice analysis algorithms (e.g., OpenCV, speech_recognition) to analyze the user's facial expressions and tone of voice. The input is facial image data and voice data, and the output is the user's emotional data (e.g., "happiness," "excitement," etc.).
[1424] Step 4:
[1425] The device generates emotion data and transmits it to the server along with the captured data using a secure communication protocol (e.g., HTTPS). The input is facial image data, voice data, and emotion data, and the output is the securely transmitted data.
[1426] Step 5:
[1427] The server inputs the received facial image data, voice data, and emotion data into a generative AI model. The generative AI model then replaces the faces and voices of characters in the content with the user's face and voice. The input is the captured data and emotion data, and the output is the converted face and voice of the characters.
[1428] Step 6:
[1429] The server adjusts the facial expressions and tone of voice of the characters in the content based on the emotional data. The generative AI model changes the facial expressions and voice of the characters in real time to match the user's emotions. The input is emotional data, and the output is adjusted facial expression data and tone of voice data.
[1430] Step 7:
[1431] The server analyzes the dialogue of other characters and converts it into new dialogue that includes the user's name. It generates the dialogue using a natural language processing algorithm (e.g., GPT-3). The input is the existing dialogue data, and the output is the new dialogue data.
[1432] Step 8:
[1433] The server regenerates the content based on the facial data, voice data, adjusted facial expression data, and new dialogue data. This data is then integrated using video editing software (e.g., Adobe Premiere Pro). The input is all the converted data, and the output is the regenerated content.
[1434] Step 9:
[1435] The server compresses and encodes the regenerated content, encoding it at the appropriate bitrate, and sends the compressed data to the device using a secure communication protocol. The input is the regenerated content, and the output is the compressed and encoded content.
[1436] Step 10:
[1437] The device receives the regenerated content and saves it to local storage. It then launches the device's media player and begins playing the new content. The input is the compressed and encoded content, and the output is a customized video played to the user.
[1438] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1439] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1440] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1441] [Fourth embodiment]
[1442] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1443] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1444] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1445] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1446] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1447] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1448] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1449] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1450] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1451] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1452] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1453] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1454] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1455] The present invention is a system for personalizing the viewing experience of movies and video content. Users can use their own face and voice to replace characters in the content with themselves. This system is characterized by exchanging data using a secure communication protocol and playing the generated customized content on the user's terminal. Specific embodiments of the system are described below.
[1456] User operations
[1457] 1. Launching the application
[1458] The user launches an application on the device by tapping the app icon on the device's home screen.
[1459] 2. Face and voice capture
[1460] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1461] When a user faces the camera and speaks into the microphone for a few seconds, the device captures and records video and audio in real time.
[1462] 3. Content Selection
[1463] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[1464] Server Processing
[1465] 1. Data Receipt and Analysis
[1466] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a voice synthesis model.
[1467] 2. Acquiring character data
[1468] The server retrieves character data (face, voice, movements, lines) from the database of the selected content.
[1469] 3. Use of generative AI
[1470] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the inputted face and voice data of the user.
[1471] 4. Generating new lines
[1472] The server changes the lines that other characters use to call out to the user to new lines that include the user's name, so that other characters will call out the user's name.
[1473] 5. Building regenerative content
[1474] The server reconstructs the original video content based on the generated facial, voice, and dialogue data, and then uses editing software to integrate the new characters and voices into the original data.
[1475] 6. Data Compression and Transfer
[1476] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[1477] Terminal handling
[1478] 1. Receiving new content
[1479] The terminal receives the regenerated content sent from the server and stores it in local storage.
[1480] 2. Playing content
[1481] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1482] Specific examples
[1483] For example, suppose Person A captures his or her face and voice and selects an animated movie. The server inputs Person A's face and voice into the generation AI, which replaces the face and voice of the main character in the animated movie with Person A's. It also changes the lines of other characters to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can play the content in an application and enjoy the animated movie starring himself or herself as the main character.
[1484] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent the generation of inappropriate content.
[1485] The processing flow will be explained below.
[1486] Step 1:
[1487] The user launches an application on the device by tapping the corresponding app icon on the device's home screen.
[1488] Step 2:
[1489] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1490] Step 3:
[1491] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[1492] Step 4:
[1493] The user selects the movie or video content they want to watch through the device's interface, and the device selects the title from its content library.
[1494] Step 5:
[1495] The device transmits the captured face and voice data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a voice synthesis model.
[1496] Step 6:
[1497] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1498] Step 7:
[1499] The server uses generative AI to replace the face and voice of the characters in the content with those of the user based on the inputted face and voice data of the user, and generates a digital model that reflects the user's characteristics.
[1500] Step 8:
[1501] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1502] Step 9:
[1503] The server recreates the original video content based on the generated facial, voice, and dialogue data, and uses editing software to integrate the new characters and voices into the original data.
[1504] Step 10:
[1505] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[1506] Step 11:
[1507] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[1508] Step 12:
[1509] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1510] Example 1
[1511] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] In content such as videos and movies, it is desirable for users to replace characters with themselves to have a more personal and immersive experience. However, this process is technically complex, and there has been no system that makes it easy for users to do so. In addition, security measures to prevent inappropriate content generation are often lacking, so a means of safely providing customized content has been sought.
[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1514] In this invention, the server includes means for transmitting the captured face and voice data to the server using a secure communication protocol, means for replacing the faces and voices of characters in the content with the face and voice of the user based on the input data, and means for changing the lines of other characters to new lines including the user's name, thereby enabling the user to safely and easily replace themselves with characters in the content and enjoy a personal and immersive experience.
[1515] "Means for capturing a user's face and voice in real time" refers to a system that uses a camera and a microphone to record and record a user's facial expressions and voice in real time.
[1516] "Means for inputting captured facial and audio data into generative AI" refers to the process or interface for inputting facial images and audio data captured in real time into an AI model.
[1517] "Means of replacing the face and voice of a character in content with the face and voice of the user" refers to technology that converts and replaces the face and voice of a specific character in content based on acquired face and voice data of the user.
[1518] The "means for changing the lines of other characters to new lines that include the user's name" is a text processing algorithm for modifying the lines spoken by other characters to include the user's name.
[1519] "Means for regenerating content" refers to methods and tools for editing and recomposing the original content based on the replaced face and voice data.
[1520] "Means for compressing and encoding the reproduced content" means techniques for reducing the size of the reproduced video and audio data and converting it into an appropriate format.
[1521] The "means for transmitting to the user's terminal" refers to a communication means for safely and quickly transmitting data generated from the server side to the user's device.
[1522] The "means for the user's terminal to play the regenerated content" refers to software or playback functions that allow the user's device to properly play the acquired customized content.
[1523] The present invention is a system that allows users to replace themselves with characters in movies and video content to enjoy a personalized viewing experience. The system includes a means for capturing the user's face and voice in real time and using a generative AI model to replace the face and voice of characters in the content with that of the user.
[1524] First, the user launches the application installed on their device, which can be a smartphone or tablet. The user grants camera and microphone access permissions from the application's main screen and captures their face and voice. This captured data is encoded in real time and sent to the server using a secure communication protocol (e.g., HTTPS).
[1525] The server uses a facial recognition algorithm (e.g., OpenCV) and a voice synthesis model (e.g., WaveNet) to analyze the received face and voice data. The server simultaneously retrieves the face, voice, movement, and dialogue data of the character from a database of the selected content. A generative AI model (e.g., DeepFake or DALL-E) is then used to replace the character's face and voice with the user's.
[1526] Additionally, a text processing algorithm is run to replace dialogue spoken by other characters with new dialogue that includes the user's name, causing other characters to call out the user's name. The server then uses this processed data to reconstruct the original video content in editing software such as Adobe Premiere Pro or Final Cut Pro. During this editing process, the new faces, voices, and dialogue are integrated into the timeline to create a seamless visual experience.
[1527] The regenerated content is compressed and encoded (e.g., H.264 codec) to optimize file size, and then transmitted to the device using a secure communication protocol. The user can then play the regenerated content stored in the device's local storage using the application's media player and enjoy customized movies and videos featuring themselves.
[1528] As a concrete example, suppose Person A captures his or her face and voice using an app and selects his or her favorite animated movie. The system inputs Person A's face and voice into a generative AI model, which then replaces the face and voice of the main character in the animated movie with Person A's. The dialogue of other characters in the anime is also changed to call Person A's name. The newly regenerated animated movie is compressed, encoded, and sent to Person A's device. Person A can then play the content in the application and enjoy the animated movie in which he or she appears as the main character.
[1529] Example prompt sentence:
[1530] "Mr. A, please face the camera to capture your face and speak into the microphone. Then, choose your favorite movie and press the start button."
[1531] This system not only allows users to have a more personal and immersive experience, but also provides security to prevent inappropriate content from being generated.
[1532] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1533] Step 1:
[1534] Input: A user launches an application by tapping the app icon on the device's home screen.
[1535] Processing: The user launches the application by tapping the app icon on the device. The app displays a dialog requesting permission to use the camera and microphone.
[1536] Output: The application launches and displays a screen requesting camera and microphone access permissions.
[1537] What happens: The user taps the button to allow access permissions, granting access to the camera and microphone.
[1538] Step 2:
[1539] Input: The user faces the camera and speaks into the microphone for a few seconds.
[1540] Processing: The device captures the user's face and voice in real time and encodes it.
[1541] Output: The captured face and audio data is temporarily stored on the device.
[1542] What happens: The user looks at the camera and speaks into the microphone as instructed, and the application records and films this.
[1543] Step 3:
[1544] Input: Captured face and audio data.
[1545] Processing: The device sends the captured data to the server using a secure communication protocol (e.g., HTTPS).
[1546] Output: Face and voice data received by the server.
[1547] What it does: The application encrypts the captured data and sends it over the internet to a server.
[1548] Step 4:
[1549] Input: Face and voice data received by the server.
[1550] Processing: The server analyzes the data using facial recognition algorithms (e.g., OpenCV) and speech synthesis models (e.g., WaveNet).
[1551] Output: Face and voice data analyzed on the server.
[1552] Specific operation: The server analyzes facial features and voice waveforms to extract the necessary data.
[1553] Step 5:
[1554] Input: Analyzed face and audio data, selected content.
[1555] Processing: The server retrieves the character's face, voice, movement, and dialogue data from the database of the selected content.
[1556] Output: Character face data, voice data, movement data, and dialogue data.
[1557] What happens: The server queries the database for the specified content and retrieves the required data.
[1558] Step 6:
[1559] Input: Captured character data and parsed user data.
[1560] Processing: The server uses a generative AI model (e.g., DeepFake or DALL-E) to replace the character's face and voice with the user's face and voice.
[1561] Output: The replaced data.
[1562] Specific operation: The server runs the AI model and synthesizes the characters' faces and voices with those of the user.
[1563] Step 7:
[1564] Input: Replaced face and voice data and original description data.
[1565] What happens: The server changes the dialogue of the other characters to new dialogue that includes the user's name.
[1566] Output: The corrected dialogue data.
[1567] What it does: It uses a text processing algorithm to change the dialogue of characters to match the user's name.
[1568] Step 8:
[1569] Input: Replaced face and voice data and modified dialogue data.
[1570] Processing: The server uses editing software such as Adobe Premiere Pro or Final Cut Pro to reconstruct the content.
[1571] Output: The regenerated content.
[1572] How it works: The server synthesizes new data on the editing software's timeline to create a seamless video.
[1573] Step 9:
[1574] Input: The regenerated content.
[1575] Processing: The server compresses the regenerated content and encodes it with the H.264 codec.
[1576] Output: Compressed and encoded content data.
[1577] Specific operation: The server uses encoding software to optimize the file size and sends it to the device using a secure communication protocol.
[1578] Step 10:
[1579] Input: Compressed and encoded content data sent to the terminal.
[1580] Processing: The device saves the received content in local storage.
[1581] Output: Content saved to local storage.
[1582] Specific operation: The device receives the data and automatically saves it in the specified folder.
[1583] Step 11:
[1584] Input: Content stored in local storage.
[1585] Action: The device launches a media player and plays the stored content.
[1586] Output: The customized content that is played.
[1587] What happens: The user taps the play button in the app and starts watching the customized movie or video.
[1588] (Application example 1)
[1589] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1590] Traditional movie and video content viewing experiences lack personalization and immersion for viewers. It is difficult for viewers to customize the experience by replacing characters with their own faces and voices, providing a limited experience. Furthermore, there are insufficient means to securely and efficiently transmit this customized content.
[1591] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1592] In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into a generative AI model, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for regenerating the content based on the replaced data, means for transmitting the regenerated content to the user's terminal, means for the user to select content they wish to view, and means for saving the content in local storage and playing it back. This allows the user to securely and efficiently receive customized content in which they themselves are replaced with characters, and enjoy a highly immersive viewing experience.
[1593] "User" refers to the person who interacts with the system, captures their face and voice, and views customized content.
[1594] "Real-time face and voice capture method" refers to a device or software method for capturing a user's face and voice and collecting that data.
[1595] "Generative AI model" refers to an artificial intelligence model that uses captured facial and voice data to generate the face and voice of a new character.
[1596] "Input means" refers to the mechanism or method for inputting captured facial and voice data into a generative AI model.
[1597] "Characters" refers to characters that appear in movies and video content.
[1598] "Means of replacement" refers to the operation of using a generative AI model to replace the faces and voices of characters in the content with the user's face and voice.
[1599] "Regeneration means" refers to the method or process for creating new edited content based on replaced data.
[1600] "Means for playing content" refers to the mechanism or method for playing the created customized content on the user's terminal.
[1601] "Secure communication protocol" refers to a communication method for securely transmitting captured data to a server.
[1602] "Compression / Encoding Methods" means the methods and techniques that convert the data in the Reproduced Content into a format that can be efficiently transferred.
[1603] "Local storage" refers to the data storage area built into the user's device.
[1604] MODE FOR CARRYING OUT THE INVENTION
[1605] System program and processing overview
[1606] An embodiment of the present invention is a system that captures a user's face and voice to provide customized content. This system operates in conjunction with a terminal, such as a smartphone, smart glasses, a head-mounted display, or a robot, and a server. The specific operation of the system and its implementation method are described below.
[1607] Hardware and software configuration:
[1608] 1. Device:
[1609] Camera: Captures the user's face using OpenCV.
[1610] Microphone: Uses a SoundDevice to capture the user's voice.
[1611] Storage: Save data to local storage.
[1612] Communication: Use Requests to transmit data securely.
[1613] Media Player: Use MoviePy to play customized content.
[1614] 2. Server:
[1615] Generative AI model: Replaces characters in content based on the user's facial and voice data.
[1616] Database: Stores data on movies and video content.
[1617] Communication protocol: Use a communication protocol to securely transmit data.
[1618] Encoding software: compresses and encodes the reproduced content and sends it to the user's device.
[1619] Data processing and calculation flow:
[1620] 1. On the user's device:
[1621] The user launches an application on the device to capture their face and voice. The captured face image and voice data are temporarily stored in local storage.
[1622] The user selects the movie or video content they want to watch, and the content selection information is also stored in local storage.
[1623] 2. Server:
[1624] The facial and voice data sent from the device is received and input into the generative AI model, and this process uses a secure communication protocol to protect the data.
[1625] The generative AI model replaces the faces and voices of characters in the content with those of the user based on the input data.
[1626] Once all processing is complete, the regenerated content is compressed and encoded using encoding software and sent to the user's device.
[1627] 3. Playing content:
[1628] The user's device receives the new content and saves it to local storage, where the user can play the saved customized content using MoviePy.
[1629] Examples:
[1630] For example, if a user captures their face and voice and selects a particular movie, they can send the data to a server using a prompt like this:
[1631] bash
[1632] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[1633] This prompt example uploads the files face.jpg and audio.wav to the server and specifies the selected movie ID (example_movie_id). The server generates customized content based on the received data and sends it to the user's device.
[1634] As described above, by using this system, users can have a personalized experience by replacing themselves with characters in movies and video content based on their facial and voice data.
[1635] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1636] Step 1:
[1637] The user launches an application on the device. The user opens the application by tapping the app icon on the device's home screen. The user then proceeds to the face and voice capture screen.
[1638] Input: Face and voice capture request
[1639] Output: Face image (face.jpg) and audio data (audio.wav)
[1640] Specific behavior:
[1641] The device will request permission to access the camera and microphone and display a face and audio capture page. When the user faces the camera and speaks into the microphone, the device will capture video and audio in real time and save them to local storage as "face.jpg" and "audio.wav."
[1642] Step 2:
[1643] The user selects the movie or video content they want to watch by using the application's interface to select the title from the content library.
[1644] Input: User-selected content ID
[1645] Output: Selected content ID (content_id)
[1646] Specific behavior:
[1647] The device interface displays a content selection screen to the user, and the user selects the content they want to watch by tapping on it. This selection information is saved as "content_id."
[1648] Step 3:
[1649] The device transmits the captured face and voice data and the selected content ID to a server using a secure communication protocol.
[1650] Input: Face image (face.jpg), audio data (audio.wav), content ID (content_id)
[1651] Output: Send data to the server
[1652] Specific behavior:
[1653] The terminal uses the Requests library to send data to the server using a secure communication protocol, specifically using the following prompt sentence:
[1654] bash
[1655] curl -X POST "https: / / example.com / upload" -F "face=@face.jpg" -F "audio=@audio.wav" -F "content_id=example_movie_id"
[1656] Step 4:
[1657] The server analyzes the received facial and voice data and inputs it into a generative AI model, which then replaces the face and voice of the characters in the content with those of the user based on their face and voice.
[1658] Input: Face image (face.jpg), audio data (audio.wav)
[1659] Output: Replaced face and voice data
[1660] Specific behavior:
[1661] The server uses a generative AI model to analyze the user's facial image and voice data and process it to replace them with a character in the content, converting the character's face and voice to that of the user.
[1662] Step 5:
[1663] The server regenerates the content based on the replaced face and voice data, compresses and encodes the regenerated content, and sends it to the user's device.
[1664] Input: Replaced face and voice data
[1665] Output: Regenerated content files
[1666] Specific behavior:
[1667] The server uses encoding software to compress and encode the reproduced content into an efficiently transferable format, and the encoded content file is sent to the user's device using a secure communications protocol.
[1668] Step 6:
[1669] The device receives the regenerated content and stores it in local storage, allowing the user to play and enjoy the content.
[1670] Input: Regenerated content file
[1671] Output: Playable content file
[1672] Specific behavior:
[1673] The device receives the regenerated content sent from the server and stores it in local storage. The user can then use MoviePy to play the saved customized content and enjoy watching it.
[1674] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1675] The present invention provides a more interactive and emotional experience by combining a system for personalizing the viewing experience of movies and video content with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[1676] User operations
[1677] 1. Launching the application
[1678] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[1679] 2. Face and voice capture
[1680] The application asks the user to capture their face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1681] When a user faces the camera and speaks into the microphone, the device captures video and audio in real time and records them as digital data.
[1682] 3. Analysis by Emotion Engine
[1683] The captured facial and voice data is input into the emotion engine, which analyzes the user's emotions. The device sends the user's facial expressions and tone of voice to the emotion engine, which then generates emotion data.
[1684] 4. Content Selection
[1685] Users select the movie or video content they want to view through the device's interface, then select the title from the device's content library.
[1686] Server Processing
[1687] 1. Data Receipt and Analysis
[1688] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[1689] 2. Acquiring character data
[1690] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1691] 3. Use of generative AI
[1692] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[1693] 4. Generating new lines
[1694] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1695] 5. Building regenerative content
[1696] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[1697] 6. Data Compression and Transfer
[1698] The server compresses the reproduced content, encodes it at the appropriate bitrate, and sends it to the device.
[1699] Terminal handling
[1700] 1. Receiving new content
[1701] The terminal receives the regenerated content sent from the server and stores it in local storage.
[1702] 2. Playing content
[1703] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1704] Specific examples
[1705] For example, suppose Person B captures his or her face and voice, and the emotion engine detects "joy." If Person B selects an action movie, the server uses Person B's emotional data to adjust the facial expression and tone of voice of the protagonist in the action scene to match joy. Furthermore, the server changes the dialogue of other characters to call Person B's name, generating a new action movie that reflects the emotional expression. This new content is compressed, encoded, and sent to Person B's device. Person B can then play this customized action movie in the application and enjoy a more emotionally resonant viewing experience.
[1706] This system not only allows users to have a more interactive and emotional experience, but also provides security to prevent the generation of inappropriate content.
[1707] The processing flow will be explained below.
[1708] Step 1:
[1709] The user launches an application on the device. The user opens the application by tapping the corresponding app icon on the device's home screen.
[1710] Step 2:
[1711] The device asks to capture the user's face and voice. The device requests permission to access the camera and microphone and displays the face and voice capture page.
[1712] Step 3:
[1713] The user faces the camera and speaks into the microphone, and the device captures video and audio in real time and records them as digital data.
[1714] Step 4:
[1715] The device sends the captured data to the emotion engine, which analyzes the user's facial expressions and tone of voice to generate emotion data.
[1716] Step 5:
[1717] The device displays a content selection interface, allowing the user to select the movie or video content they want to view.
[1718] Step 6:
[1719] The device transmits the captured facial, voice, and emotion data using a secure communication protocol to a server, which receives it and inputs it into a facial recognition algorithm and a speech synthesis model.
[1720] Step 7:
[1721] The server retrieves character data (face, voice, movements, lines) from the database of the selected content, and then analyzes the characters' expressions and movements in detail.
[1722] Step 8:
[1723] The server uses a generative AI to replace the faces and voices of characters in the content with those of the user based on the input face and voice data of the user, and also adjusts the characters' facial expressions and tone of voice based on emotional data.
[1724] Step 9:
[1725] The server retrieves the dialogue list of other characters. The server parses the dialogue database and replaces it with a new dialogue that includes the user's name.
[1726] Step 10:
[1727] The server recreates the original video content based on the generated facial, voice, emotional, and dialogue data, and uses editing software to integrate new characters, voices, and emotional expressions into the original data.
[1728] Step 11:
[1729] The server compresses and encodes the regenerated content, encodes it at the appropriate bitrate, and sends it to the device.
[1730] Step 12:
[1731] The device receives the regenerated content sent from the server. The device starts the data reception process and stores the new video content in local storage.
[1732] Step 13:
[1733] The device plays the received content to the user, launching the device's media player and starting playback of the new content.
[1734] Example 2
[1735] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1736] Conventional movie and video content has primarily been a one-way viewing experience, making it difficult to customize it to reflect the user's emotions and individual characteristics. The generation of content that interactively responds to the viewer's emotions has been extremely limited. Furthermore, conventional technology has been inadequate for reflecting the user's face and voice onto characters in movies and videos in real time. By solving these issues, it was necessary to provide a more personal and emotional viewing experience.
[1737] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing the facial expressions and movements of characters in detail from a database, a means for analyzing the line database to obtain line lists of other characters and generate new lines, and a means for replacing the face and voice of a character based on the face and voice data of a user using the generated AI model. This makes it possible to replace the face and voice of a character in selected content with the face and voice of the user, thereby interactively customizing the content according to the user's emotions.
[1738] "User" means an individual who utilizes the System to customize movie and video content.
[1739] "Device" means a device that uses a camera and microphone to capture face and voice, including a smartphone, tablet, or computer.
[1740] "Capture" refers to the process of using a camera and microphone to digitally record facial and audio data.
[1741] "Emotion engine" refers to software that analyzes a user's emotions from captured facial and voice data and generates emotional data.
[1742] "Generative AI model" refers to an artificial intelligence model that replaces the faces and voices of characters in movies and video content with those of the user based on the user's facial and voice data.
[1743] A "dialogue database" refers to a database that stores dialogue used in movies and video content and keeps it in an analyzable format.
[1744] "Secure communication protocol" refers to the communication protocol for securely transmitting captured data to a server.
[1745] "Regenerated Content" refers to customized film or video content generated based on a user's face and voice, emotional data, and new dialogue data.
[1746] "Editing Software" means video editing tools used to create Regenerated Content, including, but not limited to, Adobe Premiere Pro.
[1747] "Compression / encoding" refers to the process of converting digital content to a specific bit rate and format to reduce the data size.
[1748] "Local storage" refers to a memory area that exists within a device and is used to store data.
[1749] The present invention provides a system that provides a more interactive and emotional experience by customizing the viewing experience of movies and video content and combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[1750] Hardware and software used
[1751] This system uses devices such as smartphones, tablets, and PCs. Each device must be equipped with a camera and microphone. The server is a back-end server that receives data using a secure communication protocol (e.g., HTTPS).
[1752] Additionally, the following specific software is used:
[1753] Emotion engine: Emotion recognition algorithm (e.g. Microsoft Azure Emotion API)
[1754] Generative AI models: Generative AI models (e.g., OpenAI GPT-3, DeepFaceLab)
[1755] Editing software: Video editing software (e.g. Adobe Premiere Pro)
[1756] Specific implementation methods
[1757] 1. Launching the application
[1758] The user launches the application by tapping the app icon on the device's home screen, and the application displays its initial screen.
[1759] 2. Face and voice capture
[1760] 1. The device will request permission to access the camera and microphone. If the user allows it, the face and voice capture page will open.
[1761] 2. The user faces the camera and speaks as instructed. The device captures this in real time and generates digital data.
[1762] 3. Analysis by Emotion Engine
[1763] The device sends the captured face and voice data to the emotion engine, which analyzes the data and generates emotion data.
[1764] 4. Content Selection
[1765] The user selects their favorite movie from the "Choose a Movie" interface. Once the user makes a selection, the device displays detailed information from the content library.
[1766] 5. Data processing by the server
[1767] 1. The device sends the capture data and emotion data to the server, which then retrieves the character data for the selected movie from the database.
[1768] 2. The server uses a generative AI model to replace the character based on the user's facial and voice data, adjusting facial expressions and tone of voice based on emotional data.
[1769] 3. The server analyzes the dialogue database and generates new dialogue.
[1770] 4. The server uses editing software such as Adobe Premiere Pro to integrate the new characters, voices, and emotional expressions into the original data.
[1771] 6. Content Compression and Transfer
[1772] The server compresses the regenerated content using an appropriate video codec, such as H.264, and sends it to the device.
[1773] 7. Receiving and Playing Content
[1774] 1. The device receives the regenerated content from the server and stores it in local storage.
[1775] 2. The user selects new content within the application and begins playback.
[1776] Specific examples
[1777] For example, if a user selects an action movie and captures their face and voice, the emotion engine may recognize the user's emotion as "joy." In this case, the server will replace the main character of the movie with the user's own character based on the user's face and voice data, and adjust the characters' facial expressions and tone of voice to reflect "joy." It will also change the dialogue so that other characters call the user's name, creating a customized action movie. This new content is then sent to the user's device, providing the user with an interactive and emotional viewing experience.
[1778] Prompt Sentence Examples
[1779] Example prompts for generative AI models:
[1780] "Based on the user's facial and voice data, replace the face and voice of the main character in an action movie with the user's. Also, if the emotion data is 'joy', adjust the main character's facial expression and tone of voice to be happy, and generate new lines for other characters in the movie to call the user's name."
[1781] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1782] Step 1:
[1783] The user launches an application on the device.
[1784] Specific operation: The user taps the app icon on the device's home screen and the application launches.
[1785] Input: User taps.
[1786] Output: The initial screen of the application is displayed.
[1787] Step 2:
[1788] The device will request permission to access the camera and microphone.
[1789] Specific behavior: The application displays "Please allow access to your camera and microphone." The user taps "Allow."
[1790] Input: User's access permission permission operation.
[1791] Output: The camera and microphone will be activated and the Face and Voice Capture page will open.
[1792] Step 3:
[1793] The user faces the camera and speaks into the microphone, which the device captures as digital data in real time.
[1794] Specific operation: The user follows the instructions to face the camera and say "hello." The device captures the facial image and audio and records them as digital data.
[1795] Input: User's face and voice.
[1796] Output: Digital data of the captured face and voice.
[1797] Step 4:
[1798] The device sends the captured face and voice data to the emotion engine.
[1799] How it works: The device uploads face and voice data to the cloud-based emotion engine, which then analyzes the data and generates emotion data.
[1800] Input: Captured digital face and voice data.
[1801] Output: Parsed emotion data.
[1802] Step 5:
[1803] The user selects the movie or video content they want to view on the device interface.
[1804] What happens: The user selects a movie from the content library and confirms the selection.
[1805] Input: User's movie title selection.
[1806] Output: Detailed information about the selected movie is displayed.
[1807] Step 6:
[1808] The device transmits the capture data and emotion data to the server.
[1809] Specific operation: The device encrypts the data and sends it to the server via a secure communication protocol.
[1810] Input: Capture data, emotion data.
[1811] Output: The server receives face, voice, and emotion data.
[1812] Step 7:
[1813] The server retrieves character data from the database and analyzes it.
[1814] Specific operation: The server accesses a movie database to obtain data on the characters' faces, voices, movements, and lines. The obtained data is then analyzed using a facial recognition algorithm and a voice synthesis model.
[1815] Input: A content database containing character data.
[1816] Output: Analyzed character face, voice, movement, and dialogue data.
[1817] Step 8:
[1818] The server uses a generative AI model to replace characters based on the user's facial and voice data.
[1819] How it works: The server uses a generative AI model to replace the user's face and voice data with the character's face and voice, and also adjusts facial expressions and tone of voice based on emotional data.
[1820] Input: Capture data, emotion data, character data.
[1821] Output: Character data replaced with the user's face and voice.
[1822] Step 9:
[1823] The server analyzes the dialogue database and generates new dialogue.
[1824] How it works: The server accesses the dialogue database and generates new dialogue for other characters to call the user's name. The dialogue generation uses a generative AI model.
[1825] Input: Dialogue database, user's name.
[1826] Output: New dialogue data containing the user's name.
[1827] Step 10:
[1828] The server recreates the original video content.
[1829] Specific operation: The server uses editing software such as Adobe Premiere Pro to integrate the generated face, voice, emotion data, and new dialogue data into the original video and re-edit it.
[1830] Input: Generated face, voice, and emotion data, new dialogue data, and original video content.
[1831] Output: The regenerated video content.
[1832] Step 11:
[1833] The server compresses and encodes the regenerated content and sends it to the device.
[1834] Specific operation: The server compresses the reproduced content using a video codec such as H.264 and sends it to the terminal via a secure communication protocol.
[1835] Input: The regenerated content.
[1836] Output: Compressed and encoded regenerated content.
[1837] Step 12:
[1838] The device receives the regenerated content and stores it in local storage.
[1839] Specific operation: The device receives the regenerated content from the server through a secure communication protocol and stores it in local storage.
[1840] Input: Compressed and encoded regenerated content.
[1841] Output: Regenerated content saved to local storage.
[1842] Step 13:
[1843] The user starts playing the regenerated content within the application on the terminal.
[1844] Specific operation: The user operates the application, selects the regenerated content, and presses the play button. The device's media player starts and the regenerated content is played.
[1845] Input: User playback operations.
[1846] Output: Playback of regenerated content.
[1847] The above processing steps provide the user with a customized and interactive viewing experience.
[1848] (Application example 2)
[1849] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1850] Traditional movie and video content could only provide fixed content and could not be customized according to the viewer's emotions or individual characteristics. This made it difficult to provide a more interactive and emotional viewing experience. Furthermore, it was not possible to reflect the viewer's name or individual characteristics, which resulted in an insufficient personalized experience.
[1851] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's face and voice in real time, means for inputting the captured face and voice data into an emotion engine and analyzing it, means for replacing the face and voice of a character in the content with the user's face and voice based on the input data, means for adjusting the character's facial expression and tone of voice based on emotion data, means for changing the lines of other characters to new lines including the user's name, and means for regenerating content based on the replaced data, the adjusted data, and the changed lines and transmitting the regenerated content to the user's terminal. This makes it possible to provide an interactive and emotional viewing experience customized for each user.
[1852] "Means for capturing the user's face and voice in real time" refers to devices or software that instantly record the facial image and voice of the user as they speak to the terminal.
[1853] "Emotion engine" refers to algorithms and software that analyze captured facial and voice data to infer a user's emotional state.
[1854] "Means for replacing the face and voice of a character in content with the user's face and voice based on input data" refers to devices or software that execute a process to replace the face and voice of a character appearing in content using collected data of the user's face and voice.
[1855] "Means for adjusting the facial expressions and tone of voice of characters" refers to devices or software that change the facial expressions and voice characteristics of characters in the content based on emotional data obtained by the emotion engine.
[1856] "Means for changing the lines of other characters into new lines that include the user's name" refers to devices or software that convert the words spoken by characters included in the content into new words that include the user's name.
[1857] "Means for regenerating content" refers to devices or software that regenerate new video and audio data based on the user's face, voice, emotional data, and changed lines.
[1858] "Means for transmitting the regenerated content to the user's device" means the equipment or software that transmits the edited or generated new video and audio to the user's device using the appropriate protocol.
[1859] This invention is a system for customizing a user's viewing experience, implemented using a smartphone application. Specifically, it captures the user's face and voice, analyzes the data with an emotion engine, and uses a generative AI model to customize the faces, voices, emotions, and lines of characters in video content in real time, providing an interactive and emotional viewing experience.
[1860] Hardware and Software Use
[1861] Device: A smartphone is used. The smartphone's front camera and microphone are used to capture face and voice. The smartphone must also have a media player installed for viewing.
[1862] Server: Runs the emotion engine, generative AI model, and dialogue conversion algorithm.
[1863] Emotion engine: Analyzes captured facial and voice data to interpret user emotions in real time. Specific software used includes facial recognition algorithms (e.g., OpenCV) and speech recognition software (e.g., the speech_recognition library).
[1864] Generative AI models: Generate the faces, voices, and expressions of characters in content based on emotional and captured data. Specific examples include face-swapping algorithms and voice synthesis models (e.g., DeepFake, Tacotron).
[1865] Dialogue conversion algorithm: An algorithm to change the dialogue of other characters into new dialogue that includes the user's name. New dialogue is generated using a natural language processing model (e.g., GPT-3).
[1866] System flow
[1867] 1. User operations
[1868] A user launches a smartphone application and uses the camera and microphone to capture their face and voice.
[1869] 2. Data Analysis
[1870] The captured face and voice data is sent to the emotion engine, which analyzes the user's emotional state and formats the results as emotion data, including the user's facial expressions and tone of voice.
[1871] 3. Server Processing
[1872] The server receives the emotion data and inputs it into a generative AI model. The face and voice of the characters in the content are replaced with the user's, and the facial expressions and tone of voice are adjusted. The lines of other characters are also converted into new lines that include the user's name.
[1873] 4. Regenerating and Submitting Content
[1874] The server regenerates the video content based on the adjusted data and the new dialogue. This new content is compressed, encoded, and sent to the user's smartphone using a secure communication protocol.
[1875] 5. Playing Content
[1876] The user's smartphone receives the regenerated content and launches a media player to play it.
[1877] Adding specific examples
[1878] For example, a user opens the app and captures their face and voice. The emotion engine then detects "happiness." Based on this result, the facial expression and voice of the protagonist of the action movie selected by the user are adjusted to reflect a state of happiness. Furthermore, the dialogue of other characters is changed to call the user's name, "Takashi."
[1879] Example prompts for generative AI models
[1880] text
[1881] "User emotion analysis data": "Joy",
[1882] "Character Layer": {
[1883] "Expression": "Joy",
[1884] "Tone": "bright",
[1885] "Name": "Takashi"
[1886] },
[1887] "Dialogue Database": "Generate new dialogue"
[1888] The customized content thus generated is then sent to the user's smartphone, allowing the user to enjoy a more emotionally resonant viewing experience.
[1889] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1890] Step 1:
[1891] A user launches a smartphone application. The user taps the app icon on the device's home screen to open it. The application requests permission to access the camera and microphone and displays a face and voice capture page. The input is the user's operation, and the output is the activation of the camera and microphone.
[1892] Step 2:
[1893] The user faces the camera and speaks into the microphone. The device captures video and records audio in real time. The captured data is temporarily stored in the device's local storage. The input is the user's face and voice, and the output is facial image data and audio data.
[1894] Step 3:
[1895] The device sends the captured facial image data and voice data to the emotion engine to analyze the emotional state. The emotion engine uses facial recognition algorithms and voice analysis algorithms (e.g., OpenCV, speech_recognition) to analyze the user's facial expressions and tone of voice. The input is facial image data and voice data, and the output is the user's emotional data (e.g., "happiness," "excitement," etc.).
[1896] Step 4:
[1897] The device generates emotion data and transmits it to the server along with the captured data using a secure communication protocol (e.g., HTTPS). The input is facial image data, voice data, and emotion data, and the output is the securely transmitted data.
[1898] Step 5:
[1899] The server inputs the received facial image data, voice data, and emotion data into a generative AI model. The generative AI model then replaces the faces and voices of characters in the content with the user's face and voice. The input is the captured data and emotion data, and the output is the converted face and voice of the characters.
[1900] Step 6:
[1901] The server adjusts the facial expressions and tone of voice of the characters in the content based on the emotional data. The generative AI model changes the facial expressions and voice of the characters in real time to match the user's emotions. The input is emotional data, and the output is adjusted facial expression data and tone of voice data.
[1902] Step 7:
[1903] The server analyzes the dialogue of other characters and converts it into new dialogue that includes the user's name. It generates the dialogue using a natural language processing algorithm (e.g., GPT-3). The input is the existing dialogue data, and the output is the new dialogue data.
[1904] Step 8:
[1905] The server regenerates the content based on the facial data, voice data, adjusted facial expression data, and new dialogue data. This data is then integrated using video editing software (e.g., Adobe Premiere Pro). The input is all the converted data, and the output is the regenerated content.
[1906] Step 9:
[1907] The server compresses and encodes the regenerated content, encoding it at the appropriate bitrate, and sends the compressed data to the device using a secure communication protocol. The input is the regenerated content, and the output is the compressed and encoded content.
[1908] Step 10:
[1909] The device receives the regenerated content and saves it to local storage. It then launches the device's media player and begins playing the new content. The input is the compressed and encoded content, and the output is a customized video played to the user.
[1910] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1911] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1912] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1913] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1914] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1915] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1916] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1917] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1918] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1919] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1920] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1921] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1922] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1923] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1924] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1925] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1926] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1927] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1928] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1929] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1930] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1931] The following is further disclosed regarding the above embodiment.
[1932] (Claim 1)
[1933] a means for capturing the user's face and voice in real time;
[1934] A means to input the captured face and voice data into the generative AI;
[1935] A means for replacing the faces and voices of characters in the content with the user's face and voice based on input data;
[1936] a means for regenerating the content based on the replaced data; and
[1937] means for transmitting the regenerated content to a user's terminal;
[1938] A system including:
[1939] (Claim 2)
[1940] 10. The system of claim 1, further comprising means for transmitting the captured face and voice data to a server using a secure communication protocol.
[1941] (Claim 3)
[1942] 10. The system of claim 1, further comprising means for compressing and encoding the reproduced content and transmitting the compressed and encoded data to the user's terminal.
[1943] "Example 1"
[1944] (Claim 1)
[1945] a means for capturing the user's face and voice in real time;
[1946] A means to input the captured face and voice data into the generative AI;
[1947] A means for replacing the faces and voices of characters in the content with the user's face and voice based on input data;
[1948] A means for changing the lines of other characters to new lines that include the user's name;
[1949] a means for regenerating the content based on the replaced data and modified dialogue;
[1950] means for compressing and encoding the regenerated content and transmitting it to the user's terminal;
[1951] a means for the user's terminal to play the regenerated content;
[1952] A system including:
[1953] (Claim 2)
[1954] 10. The system of claim 1, further comprising means for transmitting the captured face and voice data to a server using a secure communication protocol.
[1955] (Claim 3)
[1956] 10. The system of claim 1, further comprising means for compressing and encoding the reproduced content and transmitting the compressed and encoded data to the user's terminal.
[1957] "Application Example 1"
[1958] (Claim 1)
[1959] a means for capturing the user's face and voice in real time;
[1960] A means of inputting the captured face and voice data into a generative AI model; and
[1961] A means for replacing the faces and voices of characters in the content with the user's face and voice based on input data;
[1962] a means for regenerating the content based on the replaced data; and
[1963] means for transmitting the regenerated content to a user's terminal;
[1964] means for a user to select content that they wish to view;
[1965] a means for storing and playing content locally;
[1966] A system including:
[1967] (Claim 2)
[1968] 10. The system of claim 1, further comprising means for transmitting the captured face and voice data to a server using a secure communication protocol.
[1969] (Claim 3)
[1970] 10. The system of claim 1, further comprising means for compressing and encoding the reproduced content and transmitting the compressed and encoded data to the user's terminal.
[1971] "Example 2: Combining Emotion Engines"
[1972] (Claim 1)
[1973] A means for launching an application on a terminal by a user's operation;
[1974] a means for using the device's camera and microphone to capture the user's face and voice;
[1975] means for recording the captured face and audio data in real time;
[1976] means for analyzing the recorded facial and voice data to recognize the user's emotions;
[1977] means for displaying content selected by the user on the terminal;
[1978] means for transmitting the captured face and voice data and the recognized emotion data to a server using a secure communication protocol;
[1979] A means for retrieving data of a character related to the selected content from a database;
[1980] A means to replace the face and voice of a character with that of the user using an AI model generated based on the user's face and voice data;
[1981] A means to adjust the characters' facial expressions and tone of voice to match the user's emotions,
[1982] A means for analyzing dialogue data in the content and generating new dialogue;
[1983] A means to regenerate content based on newly generated face, voice, emotion data, and dialogue data;
[1984] means for compressing and encoding the regenerated content and transmitting it to the user's terminal;
[1985] means for receiving and playing the regenerated content at a user's terminal;
[1986] A system including:
[1987] (Claim 2)
[1988] The system according to claim 1, further comprising a server that analyzes the facial expressions and movements of characters in detail and acquires character data from the database.
[1989] (Claim 3)
[1990] 2. The system of claim 1, further comprising a server that analyzes the dialogue database, obtains dialogue lists of other characters, and generates new dialogue.
[1991] "Application example 2 when combining emotion engines"
[1992] (Claim 1)
[1993] a means for capturing the user's face and voice in real time;
[1994] A means of inputting and analyzing the captured face and voice data into the emotion engine;
[1995] A means for replacing the faces and voices of characters in the content with the user's face and voice based on input data;
[1996] A means to adjust the characters' facial expressions and tone of voice based on emotional data,
[1997] A means for changing the lines of other characters to new lines that include the user's name;
[1998] A means of regenerating content based on the replaced and adjusted data and modified dialogue; and
[1999] means for transmitting the regenerated content to a user's terminal;
[2000] A system including:
[2001] (Claim 2)
[2002] 10. The system of claim 1, further comprising means for transmitting the captured face and voice data and emotion data to a server using a secure communication protocol.
[2003] (Claim 3)
[2004] 10. The system of claim 1, further comprising means for compressing and encoding the reproduced content and transmitting the compressed and encoded data to the user's terminal. [Explanation of symbols]
[2005] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing the user's face and voice in real time; A means to input the captured face and voice data into the generative AI; A means for replacing the faces and voices of characters in the content with the user's face and voice based on input data; a means for regenerating the content based on the replaced data; and means for transmitting the regenerated content to a user's terminal; A system including:
2. 10. The system of claim 1, further comprising means for transmitting the captured face and voice data to a server using a secure communication protocol.
3. 10. The system of claim 1, further comprising means for compressing and encoding the reproduced content and transmitting the compressed and encoded data to the user's terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A