System

The system uses a generative AI model to automatically create karaoke videos that align with song emotions, addressing the issues of synchronization and cost, thereby enhancing the karaoke experience and reducing copyright fees.

JP2026021062APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122744
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Karaoke videos are often repetitive, out of sync with the song's worldview, and expensive to produce, with copyright fees being a significant concern.

Method used

A system that uses a generative AI model to automatically generate karaoke insert videos by quantifying the emotions in lyrics, allowing for manual fine-tuning and distribution to user terminals, thereby reducing production costs and copyright fees while ensuring synchronization with the song's worldview.

Benefits of technology

The system provides visually engaging karaoke experiences by generating videos that match the song's worldview, reducing production costs and copyright fees, and enhancing user immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021062000001_ABST
    Figure 2026021062000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting music and lyrics and quantifying the emotion of the lyrics; means for using the quantified emotion as inputs and automatically generating a video using a generative AI model; means for reviewing and, if necessary, fine-tuning the generated video; means for uploading the final video to a database and distributing it to user terminals; and means for playing the video corresponding to the selected music when the user selects the music on the karaoke terminals.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Karaoke is popular worldwide, but its insert videos tend to be repetitive and often out of sync with the times and the worldview of the songs. Furthermore, video production is expensive, and music videos or live footage may not exist for older songs or songs included in albums. Furthermore, copyright fees are incurred when using existing music videos or live footage. The objective of this invention is to provide videos that are visually enjoyable for karaoke users while resolving these issues. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means: A system is provided that includes: a means for inputting music data and lyric data and quantifying the emotion of the lyrics; a means for automatically generating videos using a generative AI model that uses the quantified emotion data as input data; a means for reviewing the generated videos and fine-tuning them as needed; a means for uploading the final videos to a database and distributing them to a user's terminal; and a means for playing a video corresponding to a song selected by a user on a karaoke terminal. This allows karaoke insert videos to fit the worldview of the song and also reduces copyright fees.

[0006] "Music data" is digital data that includes audio information of music.

[0007] "Lyrics data" is digital data that includes text information about lyrics that accompany a song.

[0008] "Quantifying emotions" refers to analyzing the content of lyrics and expressing emotions such as joy, anger, sadness, and happiness numerically.

[0009] A "generative AI model" refers to a group of algorithms that use artificial intelligence techniques to generate specific outputs from input data.

[0010] "Automatic video generation" refers to the automatic creation of video content using artificial intelligence based on specific input data.

[0011] "Review" refers to the process of checking the generated video and evaluating its quality and suitability.

[0012] "Tweaking" refers to manual editing to correct inadequacies discovered through review and improve the quality of the video.

[0013] A "database" is a digital information resource that stores information systematically and makes it easy to search and use.

[0014] "User terminal" refers to a digital device that is directly operated by a user (e.g., karaoke machine, PC, smartphone, etc.).

[0015] "Inserted video" refers to video content that is displayed while a karaoke song is being played.

[0016] The "worldview of a song" refers to the visual and auditory expression that encompasses the themes, emotions, images, etc. conveyed by the song and lyrics.

[0017] "Copyright fees" refers to the amount that must be paid when using copyrighted material such as music videos and live footage. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention provides a system for automatically generating karaoke insert videos, and a specific embodiment thereof will be described below.

[0040] System Overview

[0041] The system is mainly composed of three components: a server, a device, and a user. The server automatically generates videos using a generative AI model and manages distribution to the device. The device provides the karaoke execution environment and an interface for users to enjoy karaoke. Users select a song and sing karaoke while visually enjoying the inserted video that matches the song.

[0042] Program processing overview

[0043] 1. The server collects video data from TV dramas and movies and corresponding music data, and prepares it as a training dataset for the generative AI model, which then learns the relationship between video scenes and music.

[0044] 2. The server inputs the new karaoke song and its lyrics, analyzes the emotions in the lyrics, and converts them into numerical values. For example, the main emotions in the lyrics are quantified as "joy 60%, sadness 40%."

[0045] 3. The server inputs the quantified emotional data into a generative AI model and automatically generates a video based on the emotional parameters. For example, for a song with a spring cherry blossom theme, a video including scenes of cherry blossoms in full bloom and students graduating will be generated.

[0046] 4. The server reviews the generated video and manually adjusts it as needed, for example by correcting the timing of certain scenes to optimize the synchronization between the video and the music.

[0047] 5. The server uploads the final video to a database and delivers it to the karaoke terminals, with the video data being transferred over a high-speed, stable connection.

[0048] 6. When the user selects a song, the device plays a video corresponding to that song. For example, if the user selects the song "Sakura," a pre-generated video of cherry blossoms in full bloom will be played.

[0049] 7. Users can visually enjoy the inserted video while singing the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0050] Specific examples

[0051] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0052] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0053] 2. Server: The quantified emotional data is input into the generative AI model, and a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony that match the content of the lyrics.

[0054] 3. Server: Preview the generated video and manually fine-tune the timing of scenes, further refining the synchronization between the video and music.

[0055] 4. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0056] 5. Device: When the user selects "Sakura," the video corresponding to the selected song will be played.

[0057] 6. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of graduation ceremonies.

[0058] In this way, this system can provide insert videos that match the worldview of the song when users enjoy karaoke, improving the visual entertainment value.In addition, by using a generative AI model, it is possible to significantly reduce the cost of video production and the burden of copyright fees.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[0062] Step 2:

[0063] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and irrelevant data.

[0064] Step 3:

[0065] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and letting it learn the associations between videos and music based on the dataset.

[0066] Step 4:

[0067] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[0068] Step 5:

[0069] The server analyzes the emotions in the lyrics and converts them into numerical values, such as joy, anger, sadness, and pleasure. In this case, the result might be "70% joy, 30% sadness."

[0070] Step 6:

[0071] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[0072] Step 7:

[0073] The server reviews the generated video and manually adjusts it as needed, correcting scene timing or other inappropriate parts.

[0074] Step 8:

[0075] The server uploads the final video to a database, including meta information about the video data (such as song title, lyrics, and sentiment analysis results).

[0076] Step 9:

[0077] The terminal synchronizes with the karaoke system's database to receive new video data, ensuring that data transfer is fast and stable.

[0078] Step 10:

[0079] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[0080] Step 11:

[0081] The device plays the video corresponding to the song selected by the user. While the song "Sakura" is playing, scenes of cherry blossoms in full bloom and a graduation ceremony are displayed.

[0082] Step 12:

[0083] Users can visually enjoy the inserted video while singing along to the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0084] Example 1

[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0086] In conventional karaoke systems, manually creating videos that match a song is time-consuming and costly, making it difficult to provide users with a satisfying entertainment experience. Furthermore, it is not possible to generate videos that match the emotions of the lyrics in real time, making it difficult to effectively convey the worldview of the song. To solve these issues, there is a need for a system that can efficiently and automatically generate videos that match a song and provide users with a high-quality karaoke experience.

[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0088] In this invention, the server includes means for inputting music data and text data and quantifying emotions in the text data, means for using the quantified emotional data as input data and automatically generating videos using a generative AI model, means for reviewing the generated videos and fine-tuning them as necessary, means for uploading the final videos to a database and distributing them to a user terminal, and means for playing videos corresponding to the selected song when the user selects a song on the karaoke terminal. This makes it possible to efficiently and automatically generate videos suited to songs and provide users with a high-quality karaoke experience.

[0089] "Music data" is digital data of audio information that makes up music.

[0090] "Character data" is digital data that includes text information such as lyrics.

[0091] A "means for quantifying emotions" is a method or device that analyzes the content of character data and expresses its emotional elements as numerical values.

[0092] A "generative AI model" is an algorithm or program that uses machine learning and deep learning to automatically generate images.

[0093] "Means for automatically generating video" refers to a method or device that uses quantified emotional data and utilizes a generative AI model to create video.

[0094] The "reviewing means" refers to a method or device for checking the generated video and correcting it if necessary.

[0095] A "database" is an information storage system for storing and managing video and music data.

[0096] A "user terminal" is a device or apparatus that a user uses to enjoy karaoke.

[0097] A "karaoke terminal" is a device or apparatus dedicated to playing karaoke songs and displaying images.

[0098] The "video playback means" refers to a method or device for displaying video in accordance with the song selected on the karaoke terminal.

[0099] The present invention is a system for automatically generating images suitable for karaoke songs. A specific embodiment of this system will be described below.

[0100] The system mainly consists of three components: a server, a device, and a user. The server has a high-performance processor (e.g., an NVIDIA GPU) and large storage capacity, and uses machine learning frameworks such as TensorFlow or PyTorch to train and run the generative AI model. The device is an Android or iOS device, or a dedicated karaoke machine with built-in video playback software. The user uses a smartphone or a device in a karaoke booth.

[0101] The server first collects video data from TV dramas and movies and the corresponding music data. This creates a dataset for learning the association between video scenes and music. Specifically, the server obtains video and music data from the Internet or a dedicated database, and then formats them.

[0102] Next, the server inputs the new song and its text data. The text data includes information such as lyrics and title. This is analyzed using natural language processing (NLP) tools to quantify the emotions. For example, the lyrics of the song "Sakura" are quantified as "60% joy, 40% sadness." The analysis results are input into the generative AI model as prompt sentences.

[0103] Here is an example prompt: "Please analyze the emotions from the lyrics data of the specified song. Based on the analysis results, please generate an image that matches the lyrics. The emotional data is '70% joy, 30% sadness'."

[0104] The server uses the emotion data as input and automatically generates video using a generative AI model, which then creates a video based on the emotion data, generating scenes such as cherry blossoms in full bloom or a graduation ceremony.

[0105] The generated video is previewed on the server, and manual adjustments are made as needed, such as correcting the timing of specific scenes to optimize the synchronization between the video and music. Once the final adjustments are complete, the video is uploaded to the database.

[0106] The server delivers the videos uploaded to the database to the user's device. This delivery is done via a high-speed, stable connection. When the user selects a song, the device plays the video corresponding to that song in real time. For example, if the user selects "Sakura," a video of cherry blossoms in full bloom will be played as the intro begins.

[0107] Users can visually enjoy the video while singing the song. Because the video scenes are linked to the lyrics, users can immerse themselves deeply in the world of the song while singing. For example, while singing the song "Sakura," users can visually enjoy scenes of cherry blossoms in full bloom in spring or a graduation ceremony. In this way, users can be provided with a high-quality karaoke experience.

[0108] This system significantly reduces the cost of video production while providing users with video that matches the worldview of the song.

[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0110] Step 1:

[0111] The server collects video data of TV dramas and movies and corresponding music data from the Internet or a dedicated database. This data is used as a training dataset for the generative AI model. The input data is the video and music data of TV dramas and movies, and the output data is a formatted training dataset. Specifically, the server obtains a scene from a movie and its background music and formats it into a dataset.

[0112] Step 2:

[0113] The server receives the newly added song and its lyrics data. It then uses a natural language processing (NLP) tool to quantify the emotion of the lyrics. The input data is the song and lyrics data, and the output data is the quantified emotion data. Specifically, the server analyzes the lyrics of the song "Sakura" and generates emotion data of "60% joy, 40% sadness."

[0114] Step 3:

[0115] The server converts the quantified emotional data into a prompt sentence and inputs it into the generative AI model. At this time, the generative AI model automatically generates a video based on the emotional parameters. The input data is the quantified emotional data, and the output data is the generated video. Specifically, the generated video includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0116] Step 4:

[0117] The server reviews the generated footage and manually adjusts it as necessary. The input data is the generated footage, and the output data is the adjusted footage. Specifically, the server corrects the timing of specific scenes and optimizes the synchronization between the footage and the music.

[0118] Step 5:

[0119] The server uploads the final video to a database and distributes it to the user's device. The input data is the adjusted video, and the output data is the video stored in the database and the distributed video data. Specifically, the video data is stored in cloud storage and a link to it is sent to the device.

[0120] Step 6:

[0121] When the user selects a song, the device plays the video delivered from the server in real time. The input data is the user's song selection information and the delivered video data, and the output data is the video to be played. Specifically, when the user selects "Sakura," the intro starts and a video of cherry blossoms in full bloom is played at the same time.

[0122] Step 7:

[0123] Users visually enjoy the video while singing the song. The input data are karaoke songs and video data, and the output data is the user's karaoke experience. Specifically, users can sing "Sakura" while watching a touching graduation scene, enriching their karaoke experience.

[0124] (Application example 1)

[0125] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0126] There is a demand for combining music and video content to provide new content experiences that users can enjoy even more. However, creating visuals that match music is costly and requires a lot of effort. In addition, manual video production is time-consuming, and it is difficult to optimize the synchronization between music and video during the process. This issue needs to be resolved.

[0127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0128] In this invention, the server includes means for inputting music data and lyric data and quantifying the emotion of the lyrics, means for automatically generating videos using a generative AI model using the quantified emotion data as input data, means for reviewing the generated videos and fine-tuning them as necessary, and means for playing videos corresponding to the selected songs when a user selects a song from a content distribution service. This improves the user experience, significantly reduces the cost and effort of video production, and enables highly accurate synchronization of music and video.

[0129] "Music data" refers to music information stored in digital format.

[0130] "Lyrics data" refers to the text information of the lyrics of a song stored in digital format.

[0131] "Emotional data" refers to the quantified information of emotional elements analyzed from lyrics.

[0132] A "generative AI model" refers to a machine learning model that uses artificial intelligence to generate new outputs from input data.

[0133] "Means for automatically generating videos" refers to the process of using an AI model to automatically create videos based on emotional data.

[0134] "Means of reviewing and fine-tuning as necessary" refers to the process of checking the generated video and manually correcting it as necessary.

[0135] "Content distribution service" refers to a service that provides digital content to users via the Internet.

[0136] "User terminal" refers to a device that a user uses to connect to the Internet, such as a smartphone or tablet.

[0137] This invention will explain a method for constructing a system in which music data and lyrics data are input, and related inserted animation is automatically generated and provided to the user.

[0138] System configuration

[0139] The system consists of the following hardware and software:

[0140] Server: Cloud server. Manages music data, analyzes lyrics sentiment, runs generative AI models, and manages generated videos.

[0141] Generative AI models: Use generative models such as GPT-4 and DALL-E.

[0142] Lyric analysis software: A text analysis tool that uses natural language processing technology.

[0143] Streaming software: Use HLS (HTTP Live Streaming) or similar.

[0144] User device: The device through which the user receives content, such as a smartphone or tablet.

[0145] Operational Overview

[0146] 1. The user selects a song

[0147] The user selects a song using the application of the content distribution service.

[0148] 2. The server retrieves the music data and lyrics data

[0149] Based on the song ID selected by the user, the server retrieves the song data and lyrics data.

[0150] 3. Lyrics Analysis

[0151] The server inputs the acquired lyrics data into lyrics analysis software and quantifies the emotions of the lyrics. For example, the main emotional elements of the lyrics are quantified as "joy 60% and sadness 40%."

[0152] 4. Visual Generation

[0153] The server inputs the quantified emotional data into a generative AI model to generate video visuals based on the emotions. For example, based on emotional data of "60% joy, 40% sadness," a video containing scenes of cherry blossoms in full bloom and students graduating might be generated.

[0154] 5. Review and fine-tune

[0155] The video generated on the server is previewed and manually tweaked as needed, correcting the timing of specific scenes and optimizing the synchronization of the video and music.

[0156] 6. Video distribution

[0157] The final fine-tuned video is delivered to the user's device using streaming software (such as HLS).

[0158] Specific examples

[0159] The prompt text and system behavior when the user selects the song "Spring Breeze" in the app are shown below.

[0160] Prompt statement

[0161] Visualize a spring landscape with emotional parameters of 70% hope and 30% joy.

[0162] System operation example

[0163] 1. When the user selects the song "Spring Breeze," the server retrieves the corresponding song data and lyrics data.

[0164] 2. Analyze the acquired lyrics data and quantify the emotional data (hope 70%, joy 30%).

[0165] 3. This emotional data is input into a generative AI model to generate images that express hope and joy, such as the fresh greenery of spring and bright skies.

[0166] 4. The generated video is reviewed on the server side, and the timing is fine-tuned to achieve optimal synchronization.

[0167] 5. The final video will be delivered to the user’s smartphone, where they can enjoy listening to the song “Spring Breeze.”

[0168] In this way, the synchronization between the music and the video can be improved, providing the user with a more attractive content experience.

[0169] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0170] Step 1:

[0171] The server receives an operation by the user to select a song and acquires the song ID.

[0172] Input: User-selected song ID

[0173] Output: Song ID is passed to the system

[0174] How it works: A user selects the song "Spring Breeze" on the app. The server detects the selection and retrieves the song ID (e.g., "12345").

[0175] Step 2:

[0176] The server uses the song ID to obtain the song data and lyrics data.

[0177] Input: Song ID

[0178] Output: Music data and lyrics data

[0179] Operation: The server accesses the database and retrieves the music data and lyrics data corresponding to the music ID "12345." At this time, the music data is retrieved in music file format, and the lyrics data is retrieved in text format.

[0180] Step 3:

[0181] The server analyzes the acquired lyrics data and quantifies the emotion of the lyrics.

[0182] Input: Lyrics data

[0183] Output: Emotion data (e.g., hope 70%, joy 30%)

[0184] How it works: The server inputs the lyrics data into natural language processing software, which analyzes the lyrics text. This analysis quantifies the emotional elements contained in the lyrics and generates emotional data, such as "70% hope, 30% joy."

[0185] Step 4:

[0186] The server inputs the quantified emotional data into a generative AI model and automatically generates videos based on the emotions.

[0187] Input: Emotion data

[0188] Output: Generated video

[0189] How it works: The server inputs the prompt "Please visualize a spring landscape with emotional parameters of 70% hope and 30% joy" into a generative AI model (e.g., DALL-E or GPT-4), and generates a video based on this. For example, a visual containing fresh spring greenery and a bright sky is generated.

[0190] Step 5:

[0191] The server reviews the generated video and makes any necessary adjustments.

[0192] Input: Generated video

[0193] Output: Final fine-tuned video

[0194] How it works: Preview the video generated on the server, check the timing and sequence of scenes, and if necessary, manually tweak specific scenes and timing to optimize the synchronization of music and visuals.

[0195] Step 6:

[0196] The server then delivers the final, finely tuned video to the user's device via streaming software.

[0197] Input: Final tweaked video

[0198] Output: Streaming video to user devices

[0199] How it works: The server uses HLS streaming software to encode the final video in real time and delivers it to the user's smartphone, tablet, or other device. Users can then watch the video corresponding to the song they selected on the app.

[0200] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0201] The present invention provides a system for automatically generating karaoke insert videos, which incorporates a technology that combines an emotion engine that recognizes the user's emotions. Specific embodiments of the technology are described below.

[0202] System configuration

[0203] The system mainly consists of the following components:

[0204] 1. Server

[0205] 2. User Device

[0206] 3. Emotion Engine

[0207] Program processing overview

[0208] 1. Server: Collects video and music data from TV dramas and movies and prepares it as a training dataset for the generative AI model, allowing the model to learn the relationship between video scenes and music.

[0209] 2. Server: A new karaoke song and its lyrics are input into the system, and the emotions in the lyrics are analyzed and quantified. For example, the main emotions in the lyrics are quantified as "joy 60% and sadness 40%."

[0210] 3. Server: The emotion engine recognizes the user's emotions in real time and feeds that emotion data back to the generative AI model. The emotion engine recognizes emotions from the user's facial expressions, tone of voice, body movements, etc.

[0211] 4. Server: Based on the emotional data obtained from the emotion engine and the quantified emotional data of the lyrics, the generative AI model automatically generates the optimal video. For example, if the user's emotions are "70% joy, 30% sadness," it will generate a video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0212] 5. Server: Review the generated video and manually tweak it as needed, such as correcting scene timing or inappropriate parts.

[0213] 6. Server: Uploads the final video to the database and distributes it to the karaoke terminals, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[0214] 7. Terminal: Synchronizes with the karaoke system database and receives new video data, ensuring fast and stable data transfer.

[0215] 8. User: When a song is selected from the karaoke song list, a video corresponding to the song will be played, and the video will be adjusted in real time based on the user's emotions.

[0216] 9. Device: The device plays a video corresponding to the song selected by the user, and the emotion engine analyzes emotions in real time and adjusts the content of the video as needed.

[0217] 10. Users: They can visually enjoy the inserted video while singing the song. By adjusting the video based on the user's emotions, they can more easily immerse themselves in the world of the song, improving the karaoke experience.

[0218] Specific examples

[0219] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0220] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0221] 2. Emotion Engine: While the user is singing "Sakura," the system recognizes the user's emotions from their facial expressions, tone of voice, and body movements. For example, if the user is singing happily, it will be recognized as "80% joy, 20% sadness."

[0222] 3. Server: The emotion data recognized by the emotion engine is fed back to the generation AI model, and combined with the emotion data from the lyrics, a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0223] 4. Server: Preview the generated video and manually fine-tune the timing of scenes to optimize the synchronization between the video and music.

[0224] 5. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0225] 6. Device: When the user selects "Sakura," a video adjusted in real time based on the user's emotions is played.

[0226] 7. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of a graduation ceremony. In addition, the content of the video is adjusted in real time to match the user's emotions, allowing for a more emotionally immersive experience.

[0227] In this way, the system recognizes the user's emotions and uses that emotional data to adjust the video in real time, providing a visually enjoyable karaoke experience. Furthermore, the use of generative AI models can significantly reduce video production costs and copyright fees.

[0228] The processing flow will be explained below.

[0229] Step 1:

[0230] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[0231] Step 2:

[0232] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and inappropriate data.

[0233] Step 3:

[0234] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and training the model to recognize associations between videos and music based on the dataset.

[0235] Step 4:

[0236] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[0237] Step 5:

[0238] The server analyzes the emotions in the lyrics and quantifies emotions such as joy, anger, sadness, and happiness. In this case, the result might be "70% joy, 30% sadness."

[0239] Step 6:

[0240] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[0241] Step 7:

[0242] The server reviews the generated video and manually adjusts it to correct timing or inappropriate parts of the scene.

[0243] Step 8:

[0244] The server uploads the final video to a database, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[0245] Step 9:

[0246] The terminal synchronizes with the karaoke system's database and receives new video data, ensuring fast and stable data transfer.

[0247] Step 10:

[0248] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[0249] Step 11:

[0250] The device plays a video corresponding to the song selected by the user. As the "Sakura" video plays, the emotion engine analyzes the user's emotions in real time.

[0251] Step 12:

[0252] The emotion engine recognizes the user's emotions in real time from their facial expressions, tone of voice, and body movements.

[0253] Step 13:

[0254] The server then feeds the emotion data recognized by the emotion engine back to the generative AI model, adjusting the video content in real time. For example, if the user is having fun, it adds a cheerful scene.

[0255] Step 14:

[0256] The device displays real-time updated video to the user, with the video content tailored to the user's current emotions.

[0257] Step 15:

[0258] Users can enjoy visually tailored videos while singing along, which helps users immerse themselves in the world of the song and enhances the karaoke experience.

[0259] In this way, by incorporating an emotion engine, it becomes possible to adjust the video content according to the user's real-time emotions, providing a more immersive karaoke experience.

[0260] Example 2

[0261] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0262] In conventional karaoke systems, the inserted video is fixed, making it impossible to dynamically adjust the video according to the user's emotions or singing style, resulting in a limited karaoke experience. In addition, creating fixed videos requires a great deal of effort and cost, making efficient operation difficult.

[0263] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0264] In this invention, the server includes: a means for inputting music data and lyric data and quantifying the emotion of the lyrics; a means for automatically generating videos using a generative AI model using the quantified emotion data as input data; a means for reviewing the generated videos and fine-tuning them as necessary; an emotion recognition means for analyzing the user's facial expressions, tone of voice, and body movements in real time; a means for feeding back the analyzed emotion data in real time to the generative AI model and adjusting the content of the videos; a means for uploading the final videos to a database and distributing them to a user terminal; and a means for playing a video corresponding to a selected song when the user selects a song on the karaoke terminal. This enables dynamic video adjustment according to the user's emotion, providing a richer karaoke experience. It also improves video production efficiency and reduces operational costs.

[0265] "Music data" refers to audio data of music played on a karaoke system.

[0266] "Lyric data" is character information corresponding to a song, and is text data indicating the content of the lyrics.

[0267] An "emotion quantification means" is a method or system that performs a process of analyzing lyric content and assigning a numerical value to a particular emotional category.

[0268] A "generative AI model" is an artificial intelligence model that automatically generates videos from input data such as emotional data.

[0269] "Means for automatically generating videos" refers to the process of using a generative AI model to generate videos based on input data.

[0270] "Means for reviewing videos and fine-tuning them as needed" refers to the processes and tools used to review the content of the videos produced and adjust any inappropriate parts or timing.

[0271] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions, tone of voice, body movements, etc.

[0272] "Means for analyzing in real time" refers to a method for acquiring user emotion data in real time and executing an analysis process.

[0273] The "means for adjusting the content of the video" is a process for appropriately changing the scenes and content of the video being played based on the user's real-time emotional data.

[0274] The "means for uploading to a database" refers to the process or tool for storing the generated final video in a database such as cloud storage or a local server.

[0275] "Means for distributing to user terminals" refers to the process of transferring video data stored in the database to the karaoke terminals used by users.

[0276] The "means for playing video" refers to a function or process for playing video corresponding to a song selected on a user terminal using a display device or audio device.

[0277] The present invention is a system for automatically generating videos to be inserted into a karaoke system based on user emotion recognition. Specific embodiments of the system will be described below.

[0278] System Configuration

[0279] The system mainly consists of the following components:

[0280] 1. Server

[0281] 2. User Device

[0282] 3. Emotion Engine

[0283] Operation overview

[0284] Video and music data collection

[0285] Server: The server collects video and music data from TV dramas and movies from public databases on the Internet and from licensed content providers. This data is prepared as a dataset for training the generative AI model. Specifically, a Python script is used to download the data via an API and save it in storage. This process creates a dataset capable of learning a variety of emotional expressions and scenes.

[0286] Emotional analysis of new karaoke songs

[0287] Server: New karaoke songs and lyrics are input into the system, and natural language processing (NLP) techniques are used to quantify the emotions in the lyrics. This analysis uses emotion recognition libraries (e.g., NLTK, SpaCy). For example, the lyrics of the song "Sakura" are quantified as "70% joy, 30% sadness." This numerical data is used as input data for the generative AI model.

[0288] User Emotion Recognition

[0289] Emotion engine: Analyzes the user's facial expressions, vocal tone, and body movements in real time. This is done using a facial recognition camera, microphone, and motion sensor. The recognized emotional data is sent to a server. Specifically, facial expression analysis is performed using OpenCV, vocal tone analysis using a voice analysis tool, and body movements are analyzed using a motion sensor.

[0290] Optimal video generation

[0291] Server: Real-time emotional data obtained from the emotion engine and quantified emotional data from lyrics are input into the generative AI model. The generative AI model automatically generates optimal video scenes based on this data. For example, for "Sakura," it generates scenes of cherry blossoms in full bloom and a graduation ceremony. The prompt text in this case is as follows:

[0292] Generate an optimal insert video based on the lyrics of "Sakura." The emotional analysis of the lyrics reveals "70% joy, 30% sadness." The emotions recognized from the user's facial expressions, tone of voice, and body movements are "80% joy, 20% sadness." Based on this, create an inspiring video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0293] Video review and fine-tuning

[0294] Server: The server reviews the generated video and manually corrects any inaccuracies or timing adjustments. For example, they use video editing software such as Adobe Premiere Pro to optimize the synchronization between the video and the music.

[0295] Uploading video data

[0296] Server: Uploads the final video to a database and prepares it for distribution to user devices. This data includes meta information (song title, lyrics, sentiment analysis results). The server accesses the database and stores the video data in cloud storage or on a local server.

[0297] Data synchronization to karaoke terminal

[0298] User terminal: Periodically synchronizes with the karaoke system database to receive new video data. The terminal accesses the server and downloads data based on a set schedule.

[0299] Video playback and real-time adjustments

[0300] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. The emotion engine analyzes the user's emotions in real time and adjusts the video content as needed, allowing the user to enjoy a more emotionally immersive karaoke experience.

[0301] Through these steps, dynamic video adjustments based on user emotions are possible, providing a richer karaoke experience, while also improving video production efficiency and reducing operational costs.

[0302] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0303] Step 1:

[0304] Server: Collects video data from TV dramas and movies and music data, and prepares it as a training dataset for the generative AI model.

[0305] Input: Video and music data collected from public databases and content providers on the Internet.

[0306] Data processing: Using a Python script, the database is accessed via API, video and music data is downloaded, and stored in storage.

[0307] Output: The dataset used to train the generative AI model.

[0308] Step 2:

[0309] Server: New karaoke songs and their lyrics are input into the system, and the emotions of the lyrics are analyzed and quantified.

[0310] Input: Karaoke song and its lyrics data.

[0311] Data Arithmetic: Using an emotion recognition library (e.g., NLTK, SpaCy), parse the lyrics and assign numerical values ​​to major emotion categories, e.g., "70% joy, 30% sadness."

[0312] Output: Quantified lyrics emotion data.

[0313] Step 3:

[0314] Emotion engine: Identifies the user's facial expressions, tone of voice, and body movements in real time and sends the emotional data to the server.

[0315] Input: The user's facial expressions, tone of voice, and body movements.

[0316] Data calculation: Facial expressions, voice, and movements are captured using a facial recognition camera, microphone, and motion sensor, and then analyzed using OpenCV and voice analysis tools.

[0317] Output: Real-time sentiment data.

[0318] Step 4:

[0319] Server: Based on the emotional data from the emotion engine and the quantified emotional data from the lyrics, the generative AI model automatically generates optimal video scenes.

[0320] Input: Real-time emotion data from the emotion engine and quantified lyric emotion data.

[0321] Data computation: A generative AI model (e.g., GPT-4, DALL-E) generates video scenes based on prompts.

[0322] Output: The generated video scenes.

[0323] Step 5:

[0324] Server: Review the generated video and manually fine-tune any necessary parts.

[0325] Input: Generated video scenes.

[0326] Data processing: Use video editing software (e.g. Adobe Premiere Pro) to adjust inappropriate parts and timing.

[0327] Output: Final optimized video.

[0328] Step 6:

[0329] Server: Uploads the final video to the database and distributes it to the karaoke terminals.

[0330] Input: The final optimized video and its meta information (song title, lyrics, sentiment analysis results).

[0331] Data processing: Video data is stored in cloud storage or on a local server and prepared for distribution.

[0332] Output: Video data uploaded to the database.

[0333] Step 7:

[0334] Terminal: Synchronizes with the karaoke system database and receives new video data.

[0335] Input: New video data from the database.

[0336] Data processing: Through a periodic synchronization process, the server is accessed, video data is downloaded, and stored in the internal storage.

[0337] Output: New video data saved on your device.

[0338] Step 8:

[0339] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. An emotion recognition engine adjusts the video content in real time as needed.

[0340] Input: User-selected song, real-time emotional data.

[0341] Data calculation: During video playback, the emotion recognition engine analyzes the user's emotions and adjusts the video content and timing accordingly.

[0342] Output: Real-time adjusted video playback provides a more emotionally immersive karaoke experience.

[0343] (Application example 2)

[0344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0345] Modern advertising videos are unable to adapt their content in real time to the viewer's emotions, making it difficult to provide an optimal visual and emotional experience for each individual user, which limits the effectiveness of advertising and reduces viewer engagement.

[0346] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0347] In this invention, the server includes: means for inputting music data and lyric data and quantifying the emotion of the lyrics; means for automatically generating videos using a generative AI model using the quantified emotion data as input data; means for reviewing the generated videos and fine-tuning them as necessary; means for uploading the final videos to a database and distributing them to user terminals; means for acquiring the user's facial expressions, tone of voice, and movements in real time and analyzing their emotions; and means for the generative AI model to dynamically change the video content in real time based on the emotion data. This makes it possible to generate advertising videos that adapt to the user's emotions in real time, improve viewer engagement, and maximize advertising effectiveness.

[0348] "Music data" refers to the audio information of a musical work and its associated metadata.

[0349] "Lyric data" refers to text information about lyrics corresponding to a song.

[0350] The "emotion engine" is a function that analyzes emotions in real time from the user's facial expressions, voice, movements, etc.

[0351] "Numericalized emotion data" refers to data that expresses emotions as numerical values.

[0352] A "generative AI model" is an artificial intelligence model that generates new data or content based on input data.

[0353] "Automatic animation generation means" refers to a method or device for automatically generating animation using input data.

[0354] A "database" is a system for efficiently storing, managing, and retrieving data.

[0355] A "user terminal" is an electronic device that is directly operated by a user.

[0356] "Tweaking tool" refers to a method or device for manually adjusting generated content.

[0357] "Facial expression recognition" is a technology that analyzes a user's facial expressions and identifies their emotions based on them.

[0358] "Tone of voice" refers to characteristics such as pitch, intensity, and speed of voice, and is an element that conveys the speaker's emotions and intentions.

[0359] "Real-time acquisition" refers to the acquisition and processing of data immediately, without delay.

[0360] "Dynamic change" means that a system or content changes automatically over time or in response to events.

[0361] An "advertising video" is video content created to promote a particular product or service.

[0362] "Viewer's emotions" refer to the emotions and feelings felt by a user watching video content.

[0363] The present invention relates to a system for dynamically generating and adjusting advertising videos in real time based on user emotions. Specific embodiments of the system are described below.

[0364] System configuration

[0365] The system mainly consists of the following components:

[0366] 1. Server

[0367] 2. User Device

[0368] 3. Emotion Engine

[0369] Program processing overview

[0370] Hardware and Software Configuration

[0371] Hardware:

[0372] Smartphone camera

[0373] Smartphone microphone

[0374] software:

[0375] EmotionEngine (emotion analysis engine)

[0376] Video Generator (generative AI model)

[0377] Database management systems (e.g., SQLite)

[0378] System Operation

[0379] 1. The server first inputs the music data and lyrics data and quantifies the emotion of the lyrics.

[0380] 2. Using the quantified emotional data as input data, a generative AI model is used to automatically generate videos.

[0381] 3. Review the generated video on the server side and make any necessary adjustments.

[0382] 4. Upload the final video to the database and distribute it to the user's device.

[0383] 5. While the user is watching the advertising video, the user's device will use a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements.

[0384] 6. The emotion engine analyzes the user's emotions based on the acquired data and generates quantified emotion data.

[0385] 7. The generative AI model dynamically changes video content based on emotional data captured in real time.

[0386] 8. The user device plays the adjusted video in real time.

[0387] Specific examples

[0388] For example, consider an advertising video for a "new sports car." This advertisement dynamically changes to emphasize the speed of the car when the user is excited, and to show scenes of the car traveling through beautiful scenery when the user is relaxed. Specific examples of prompt sentences are as follows:

[0389] Example prompt:

[0390] "For an ad for a new sports car, generate a video that shows a car speeding along if the user is excited, or a video that shows a car driving through beautiful scenery if the user is relaxed."

[0391] This system can adapt to user emotions in real time and improve viewer engagement. In particular, in the case of advertising videos, dynamically changing content to match user emotions can maximize advertising effectiveness.

[0392] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0393] Step 1:

[0394] The server inputs music data and lyrics data. The server then quantifies the emotions in the lyrics and generates emotional data. Specifically, it uses natural language processing to analyze the main emotional elements of the lyrics and quantifies them, such as "70% joy, 30% sadness."

[0395] Step 2:

[0396] The server uses the quantified emotional data as input and automatically generates videos using a generative AI model. Based on the input emotional data, it constructs videos from highly relevant scenes. For example, if the emotional data is "70% joy, 30% sadness," it selects appropriate scenes (such as cherry blossoms in full bloom or a graduation ceremony) and generates a video.

[0397] Step 3:

[0398] The server reviews the generated video and makes any necessary adjustments, such as manually correcting scene timing or inappropriate parts of the generated video. As a result of the review, the final video file is generated.

[0399] Step 4:

[0400] The server uploads the final video to a database and distributes it to the user's device. Specifically, the video data is stored in the database along with meta information (song title, lyrics, and sentiment analysis results) and made accessible from the client device.

[0401] Step 5:

[0402] While the user is watching the advertising video, the device uses a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements. Specifically, the device's camera captures video data and the microphone captures audio data.

[0403] Step 6:

[0404] Based on the data acquired by the emotion engine, the user's emotions are analyzed and quantified emotional data is generated. Specifically, emotional data such as "80% excited, 20% relaxed" is generated from the user's facial expressions and tone of voice.

[0405] Step 7:

[0406] The generative AI model dynamically changes the video content based on emotional data acquired in real time. For example, if the user is excited, it will switch to scenes that emphasize the speed of cars, and if the user is relaxed, it will switch to scenes with beautiful scenery.

[0407] Step 8:

[0408] The device then plays the adjusted video in real time, ultimately displaying advertising videos that match the user's emotions, improving viewer engagement and maximizing advertising effectiveness.

[0409] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0410] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0411] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0412] [Second embodiment]

[0413] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0414] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0415] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0416] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0417] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0418] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0419] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0420] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0421] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0422] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0423] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0424] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0425] The present invention provides a system for automatically generating karaoke insert videos, and a specific embodiment thereof will be described below.

[0426] System Overview

[0427] The system is mainly composed of three components: a server, a device, and a user. The server automatically generates videos using a generative AI model and manages distribution to the device. The device provides the karaoke execution environment and an interface for users to enjoy karaoke. Users select a song and sing karaoke while visually enjoying the inserted video that matches the song.

[0428] Program processing overview

[0429] 1. The server collects video data from TV dramas and movies and corresponding music data, and prepares it as a training dataset for the generative AI model, which then learns the relationship between video scenes and music.

[0430] 2. The server inputs the new karaoke song and its lyrics, analyzes the emotions in the lyrics, and converts them into numerical values. For example, the main emotions in the lyrics are quantified as "joy 60%, sadness 40%."

[0431] 3. The server inputs the quantified emotional data into a generative AI model and automatically generates a video based on the emotional parameters. For example, for a song with a spring cherry blossom theme, a video including scenes of cherry blossoms in full bloom and students graduating will be generated.

[0432] 4. The server reviews the generated video and manually adjusts it as needed, for example by correcting the timing of certain scenes to optimize the synchronization between the video and the music.

[0433] 5. The server uploads the final video to a database and delivers it to the karaoke terminals, with the video data being transferred over a high-speed, stable connection.

[0434] 6. When the user selects a song, the device plays a video corresponding to that song. For example, if the user selects the song "Sakura," a pre-generated video of cherry blossoms in full bloom will be played.

[0435] 7. Users can visually enjoy the inserted video while singing the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0436] Specific examples

[0437] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0438] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0439] 2. Server: The quantified emotional data is input into the generative AI model, and a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony that match the content of the lyrics.

[0440] 3. Server: Preview the generated video and manually fine-tune the timing of scenes, further refining the synchronization between the video and music.

[0441] 4. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0442] 5. Device: When the user selects "Sakura," the video corresponding to the selected song will be played.

[0443] 6. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of graduation ceremonies.

[0444] In this way, this system can provide insert videos that match the worldview of the song when users enjoy karaoke, improving the visual entertainment value.In addition, by using a generative AI model, it is possible to significantly reduce the cost of video production and the burden of copyright fees.

[0445] The processing flow will be explained below.

[0446] Step 1:

[0447] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[0448] Step 2:

[0449] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and irrelevant data.

[0450] Step 3:

[0451] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and letting it learn the associations between videos and music based on the dataset.

[0452] Step 4:

[0453] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[0454] Step 5:

[0455] The server analyzes the emotions in the lyrics and converts them into numerical values, such as joy, anger, sadness, and pleasure. In this case, the result might be "70% joy, 30% sadness."

[0456] Step 6:

[0457] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[0458] Step 7:

[0459] The server reviews the generated video and manually adjusts it as needed, correcting scene timing or other inappropriate parts.

[0460] Step 8:

[0461] The server uploads the final video to a database, including meta information about the video data (such as song title, lyrics, and sentiment analysis results).

[0462] Step 9:

[0463] The terminal synchronizes with the karaoke system's database to receive new video data, ensuring that data transfer is fast and stable.

[0464] Step 10:

[0465] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[0466] Step 11:

[0467] The device plays the video corresponding to the song selected by the user. While the song "Sakura" is playing, scenes of cherry blossoms in full bloom and a graduation ceremony are displayed.

[0468] Step 12:

[0469] Users can visually enjoy the inserted video while singing along to the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0470] Example 1

[0471] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0472] In conventional karaoke systems, manually creating videos that match a song is time-consuming and costly, making it difficult to provide users with a satisfying entertainment experience. Furthermore, it is not possible to generate videos that match the emotions of the lyrics in real time, making it difficult to effectively convey the worldview of the song. To solve these issues, there is a need for a system that can efficiently and automatically generate videos that match a song and provide users with a high-quality karaoke experience.

[0473] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0474] In this invention, the server includes means for inputting music data and text data and quantifying emotions in the text data, means for using the quantified emotional data as input data and automatically generating videos using a generative AI model, means for reviewing the generated videos and fine-tuning them as necessary, means for uploading the final videos to a database and distributing them to a user terminal, and means for playing videos corresponding to the selected song when the user selects a song on the karaoke terminal. This makes it possible to efficiently and automatically generate videos suited to songs and provide users with a high-quality karaoke experience.

[0475] "Music data" is digital data of audio information that makes up music.

[0476] "Character data" is digital data that includes text information such as lyrics.

[0477] A "means for quantifying emotions" is a method or device that analyzes the content of character data and expresses its emotional elements as numerical values.

[0478] A "generative AI model" is an algorithm or program that uses machine learning and deep learning to automatically generate images.

[0479] "Means for automatically generating video" refers to a method or device that uses quantified emotional data and utilizes a generative AI model to create video.

[0480] The "reviewing means" refers to a method or device for checking the generated video and correcting it if necessary.

[0481] A "database" is an information storage system for storing and managing video and music data.

[0482] A "user terminal" is a device or apparatus that a user uses to enjoy karaoke.

[0483] A "karaoke terminal" is a device or apparatus dedicated to playing karaoke songs and displaying images.

[0484] The "video playback means" refers to a method or device for displaying video in accordance with the song selected on the karaoke terminal.

[0485] The present invention is a system for automatically generating images suitable for karaoke songs. A specific embodiment of this system will be described below.

[0486] The system mainly consists of three components: a server, a device, and a user. The server has a high-performance processor (e.g., an NVIDIA GPU) and large storage capacity, and uses machine learning frameworks such as TensorFlow or PyTorch to train and run the generative AI model. The device is an Android or iOS device, or a dedicated karaoke machine with built-in video playback software. The user uses a smartphone or a device in a karaoke booth.

[0487] The server first collects video data from TV dramas and movies and the corresponding music data. This creates a dataset for learning the association between video scenes and music. Specifically, the server obtains video and music data from the Internet or a dedicated database, and then formats them.

[0488] Next, the server inputs the new song and its text data. The text data includes information such as lyrics and title. This is analyzed using natural language processing (NLP) tools to quantify the emotions. For example, the lyrics of the song "Sakura" are quantified as "60% joy, 40% sadness." The analysis results are input into the generative AI model as prompt sentences.

[0489] Here is an example prompt: "Please analyze the emotions from the lyrics data of the specified song. Based on the analysis results, please generate an image that matches the lyrics. The emotional data is '70% joy, 30% sadness'."

[0490] The server uses the emotion data as input and automatically generates video using a generative AI model, which then creates a video based on the emotion data, generating scenes such as cherry blossoms in full bloom or a graduation ceremony.

[0491] The generated video is previewed on the server, and manual adjustments are made as needed, such as correcting the timing of specific scenes to optimize the synchronization between the video and music. Once the final adjustments are complete, the video is uploaded to the database.

[0492] The server delivers the videos uploaded to the database to the user's device. This delivery is done via a high-speed, stable connection. When the user selects a song, the device plays the video corresponding to that song in real time. For example, if the user selects "Sakura," a video of cherry blossoms in full bloom will be played as the intro begins.

[0493] Users can visually enjoy the video while singing the song. Because the video scenes are linked to the lyrics, users can immerse themselves deeply in the world of the song while singing. For example, while singing the song "Sakura," users can visually enjoy scenes of cherry blossoms in full bloom in spring or a graduation ceremony. In this way, users can be provided with a high-quality karaoke experience.

[0494] This system significantly reduces the cost of video production while providing users with video that matches the worldview of the song.

[0495] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0496] Step 1:

[0497] The server collects video data of TV dramas and movies and corresponding music data from the Internet or a dedicated database. This data is used as a training dataset for the generative AI model. The input data is the video and music data of TV dramas and movies, and the output data is a formatted training dataset. Specifically, the server obtains a scene from a movie and its background music and formats it into a dataset.

[0498] Step 2:

[0499] The server receives the newly added song and its lyrics data. It then uses a natural language processing (NLP) tool to quantify the emotion of the lyrics. The input data is the song and lyrics data, and the output data is the quantified emotion data. Specifically, the server analyzes the lyrics of the song "Sakura" and generates emotion data of "60% joy, 40% sadness."

[0500] Step 3:

[0501] The server converts the quantified emotional data into a prompt sentence and inputs it into the generative AI model. At this time, the generative AI model automatically generates a video based on the emotional parameters. The input data is the quantified emotional data, and the output data is the generated video. Specifically, the generated video includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0502] Step 4:

[0503] The server reviews the generated footage and manually adjusts it as necessary. The input data is the generated footage, and the output data is the adjusted footage. Specifically, the server corrects the timing of specific scenes and optimizes the synchronization between the footage and the music.

[0504] Step 5:

[0505] The server uploads the final video to a database and distributes it to the user's device. The input data is the adjusted video, and the output data is the video stored in the database and the distributed video data. Specifically, the video data is stored in cloud storage and a link to it is sent to the device.

[0506] Step 6:

[0507] When the user selects a song, the device plays the video delivered from the server in real time. The input data is the user's song selection information and the delivered video data, and the output data is the video to be played. Specifically, when the user selects "Sakura," the intro starts and a video of cherry blossoms in full bloom is played at the same time.

[0508] Step 7:

[0509] Users visually enjoy the video while singing the song. The input data are karaoke songs and video data, and the output data is the user's karaoke experience. Specifically, users can sing "Sakura" while watching a touching graduation scene, enriching their karaoke experience.

[0510] (Application example 1)

[0511] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0512] There is a demand for combining music and video content to provide new content experiences that users can enjoy even more. However, creating visuals that match music is costly and requires a lot of effort. In addition, manual video production is time-consuming, and it is difficult to optimize the synchronization between music and video during the process. This issue needs to be resolved.

[0513] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0514] In this invention, the server includes means for inputting music data and lyric data and quantifying the emotion of the lyrics, means for automatically generating videos using a generative AI model using the quantified emotion data as input data, means for reviewing the generated videos and fine-tuning them as necessary, and means for playing videos corresponding to the selected songs when a user selects a song from a content distribution service. This improves the user experience, significantly reduces the cost and effort of video production, and enables highly accurate synchronization of music and video.

[0515] "Music data" refers to music information stored in digital format.

[0516] "Lyrics data" refers to the text information of the lyrics of a song stored in digital format.

[0517] "Emotional data" refers to the quantified information of emotional elements analyzed from lyrics.

[0518] A "generative AI model" refers to a machine learning model that uses artificial intelligence to generate new outputs from input data.

[0519] "Means for automatically generating videos" refers to the process of using an AI model to automatically create videos based on emotional data.

[0520] "Means of reviewing and fine-tuning as necessary" refers to the process of checking the generated video and manually correcting it as necessary.

[0521] "Content distribution service" refers to a service that provides digital content to users via the Internet.

[0522] "User terminal" refers to a device that a user uses to connect to the Internet, such as a smartphone or tablet.

[0523] This invention will explain a method for constructing a system in which music data and lyrics data are input, and related inserted animation is automatically generated and provided to the user.

[0524] System configuration

[0525] The system consists of the following hardware and software:

[0526] Server: Cloud server. Manages music data, analyzes lyrics sentiment, runs generative AI models, and manages generated videos.

[0527] Generative AI models: Use generative models such as GPT-4 and DALL-E.

[0528] Lyric analysis software: A text analysis tool that uses natural language processing technology.

[0529] Streaming software: Use HLS (HTTP Live Streaming) or similar.

[0530] User device: The device through which the user receives content, such as a smartphone or tablet.

[0531] Operational Overview

[0532] 1. The user selects a song

[0533] The user selects a song using the application of the content distribution service.

[0534] 2. The server retrieves the music data and lyrics data

[0535] Based on the song ID selected by the user, the server retrieves the song data and lyrics data.

[0536] 3. Lyrics Analysis

[0537] The server inputs the acquired lyrics data into lyrics analysis software and quantifies the emotions of the lyrics. For example, the main emotional elements of the lyrics are quantified as "joy 60% and sadness 40%."

[0538] 4. Visual Generation

[0539] The server inputs the quantified emotional data into a generative AI model to generate video visuals based on the emotions. For example, based on emotional data of "60% joy, 40% sadness," a video containing scenes of cherry blossoms in full bloom and students graduating might be generated.

[0540] 5. Review and fine-tune

[0541] The video generated on the server is previewed and manually tweaked as needed, correcting the timing of specific scenes and optimizing the synchronization of the video and music.

[0542] 6. Video distribution

[0543] The final fine-tuned video is delivered to the user's device using streaming software (such as HLS).

[0544] Specific examples

[0545] The prompt text and system behavior when the user selects the song "Spring Breeze" in the app are shown below.

[0546] Prompt statement

[0547] Visualize a spring landscape with emotional parameters of 70% hope and 30% joy.

[0548] System operation example

[0549] 1. When the user selects the song "Spring Breeze," the server retrieves the corresponding song data and lyrics data.

[0550] 2. Analyze the acquired lyrics data and quantify the emotional data (hope 70%, joy 30%).

[0551] 3. This emotional data is input into a generative AI model to generate images that express hope and joy, such as the fresh greenery of spring and bright skies.

[0552] 4. The generated video is reviewed on the server side, and the timing is fine-tuned to achieve optimal synchronization.

[0553] 5. The final video will be delivered to the user’s smartphone, where they can enjoy listening to the song “Spring Breeze.”

[0554] In this way, the synchronization between the music and the video can be improved, providing the user with a more attractive content experience.

[0555] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0556] Step 1:

[0557] The server receives an operation by the user to select a song and acquires the song ID.

[0558] Input: User-selected song ID

[0559] Output: Song ID is passed to the system

[0560] How it works: A user selects the song "Spring Breeze" on the app. The server detects the selection and retrieves the song ID (e.g., "12345").

[0561] Step 2:

[0562] The server uses the song ID to obtain the song data and lyrics data.

[0563] Input: Song ID

[0564] Output: Music data and lyrics data

[0565] Operation: The server accesses the database and retrieves the music data and lyrics data corresponding to the music ID "12345." At this time, the music data is retrieved in music file format, and the lyrics data is retrieved in text format.

[0566] Step 3:

[0567] The server analyzes the acquired lyrics data and quantifies the emotion of the lyrics.

[0568] Input: Lyrics data

[0569] Output: Emotion data (e.g., hope 70%, joy 30%)

[0570] How it works: The server inputs the lyrics data into natural language processing software, which analyzes the lyrics text. This analysis quantifies the emotional elements contained in the lyrics and generates emotional data, such as "hope 70%, joy 30%."

[0571] Step 4:

[0572] The server inputs the quantified emotional data into a generative AI model and automatically generates videos based on the emotions.

[0573] Input: Emotion data

[0574] Output: Generated video

[0575] How it works: The server inputs the prompt "Please visualize a spring landscape with emotional parameters of 70% hope and 30% joy" into a generative AI model (e.g., DALL-E or GPT-4), and generates a video based on this. For example, a visual containing fresh spring greenery and a bright sky is generated.

[0576] Step 5:

[0577] The server reviews the generated video and makes any necessary adjustments.

[0578] Input: Generated video

[0579] Output: Final fine-tuned video

[0580] How it works: Preview the video generated on the server, check the timing and sequence of scenes, and if necessary, manually tweak specific scenes and timing to optimize the synchronization of music and visuals.

[0581] Step 6:

[0582] The server then delivers the final, finely tuned video to the user's device via streaming software.

[0583] Input: Final tweaked video

[0584] Output: Streaming video to user devices

[0585] How it works: The server uses HLS streaming software to encode the final video in real time and delivers it to the user's smartphone, tablet, or other device. Users can then watch the video corresponding to the song they selected on the app.

[0586] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0587] The present invention provides a system for automatically generating karaoke insert videos, which incorporates a technology that combines an emotion engine that recognizes the user's emotions. Specific embodiments of the technology are described below.

[0588] System configuration

[0589] The system mainly consists of the following components:

[0590] 1. Server

[0591] 2. User Device

[0592] 3. Emotion Engine

[0593] Program processing overview

[0594] 1. Server: Collects video and music data from TV dramas and movies and prepares it as a training dataset for the generative AI model, allowing the model to learn the relationship between video scenes and music.

[0595] 2. Server: A new karaoke song and its lyrics are input into the system, and the emotions in the lyrics are analyzed and quantified. For example, the main emotions in the lyrics are quantified as "joy 60% and sadness 40%."

[0596] 3. Server: The emotion engine recognizes the user's emotions in real time and feeds that emotion data back to the generative AI model. The emotion engine recognizes emotions from the user's facial expressions, tone of voice, body movements, etc.

[0597] 4. Server: Based on the emotional data obtained from the emotion engine and the quantified emotional data of the lyrics, the generative AI model automatically generates the optimal video. For example, if the user's emotions are "70% joy, 30% sadness," it will generate a video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0598] 5. Server: Review the generated video and manually tweak it as needed, such as correcting scene timing or inappropriate parts.

[0599] 6. Server: Uploads the final video to the database and distributes it to the karaoke terminals, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[0600] 7. Terminal: Synchronizes with the karaoke system database and receives new video data, ensuring fast and stable data transfer.

[0601] 8. User: When a song is selected from the karaoke song list, a video corresponding to the song will be played, and the video will be adjusted in real time based on the user's emotions.

[0602] 9. Device: The device plays a video corresponding to the song selected by the user, and the emotion engine analyzes emotions in real time and adjusts the content of the video as needed.

[0603] 10. Users: They can visually enjoy the inserted video while singing the song. By adjusting the video based on the user's emotions, they can more easily immerse themselves in the world of the song, improving the karaoke experience.

[0604] Specific examples

[0605] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0606] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0607] 2. Emotion Engine: While the user is singing "Sakura," the system recognizes the user's emotions from their facial expressions, tone of voice, and body movements. For example, if the user is singing happily, it will be recognized as "80% joy, 20% sadness."

[0608] 3. Server: The emotion data recognized by the emotion engine is fed back to the generation AI model, and combined with the emotion data from the lyrics, a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0609] 4. Server: Preview the generated video and manually fine-tune the timing of scenes to optimize the synchronization between the video and music.

[0610] 5. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0611] 6. Device: When the user selects "Sakura," a video adjusted in real time based on the user's emotions is played.

[0612] 7. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of a graduation ceremony. In addition, the content of the video is adjusted in real time to match the user's emotions, allowing for a more emotionally immersive experience.

[0613] In this way, the system recognizes the user's emotions and uses that emotional data to adjust the video in real time, providing a visually enjoyable karaoke experience. Furthermore, the use of generative AI models can significantly reduce video production costs and copyright fees.

[0614] The processing flow will be explained below.

[0615] Step 1:

[0616] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[0617] Step 2:

[0618] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and inappropriate data.

[0619] Step 3:

[0620] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and training the model to recognize associations between videos and music based on the dataset.

[0621] Step 4:

[0622] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[0623] Step 5:

[0624] The server analyzes the emotions in the lyrics and quantifies emotions such as joy, anger, sadness, and happiness. In this case, the result might be "70% joy, 30% sadness."

[0625] Step 6:

[0626] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[0627] Step 7:

[0628] The server reviews the generated video and manually adjusts it to correct timing or inappropriate parts of the scene.

[0629] Step 8:

[0630] The server uploads the final video to a database, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[0631] Step 9:

[0632] The terminal synchronizes with the karaoke system's database and receives new video data, ensuring fast and stable data transfer.

[0633] Step 10:

[0634] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[0635] Step 11:

[0636] The device plays a video corresponding to the song selected by the user. As the "Sakura" video plays, the emotion engine analyzes the user's emotions in real time.

[0637] Step 12:

[0638] The emotion engine recognizes the user's emotions in real time from their facial expressions, tone of voice, and body movements.

[0639] Step 13:

[0640] The server then feeds the emotion data recognized by the emotion engine back to the generative AI model, adjusting the video content in real time. For example, if the user is having fun, it adds a cheerful scene.

[0641] Step 14:

[0642] The device displays real-time updated video to the user, with the video content tailored to the user's current emotions.

[0643] Step 15:

[0644] Users can enjoy visually tailored videos while singing along, which helps users immerse themselves in the world of the song and enhances the karaoke experience.

[0645] In this way, by incorporating an emotion engine, it becomes possible to adjust the video content according to the user's real-time emotions, providing a more immersive karaoke experience.

[0646] Example 2

[0647] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0648] In conventional karaoke systems, the inserted video is fixed, making it impossible to dynamically adjust the video according to the user's emotions or singing style, resulting in a limited karaoke experience. In addition, creating fixed videos requires a great deal of effort and cost, making efficient operation difficult.

[0649] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0650] In this invention, the server includes: a means for inputting music data and lyric data and quantifying the emotion of the lyrics; a means for automatically generating videos using a generative AI model using the quantified emotion data as input data; a means for reviewing the generated videos and fine-tuning them as necessary; an emotion recognition means for analyzing the user's facial expressions, tone of voice, and body movements in real time; a means for feeding back the analyzed emotion data in real time to the generative AI model and adjusting the content of the videos; a means for uploading the final videos to a database and distributing them to a user terminal; and a means for playing a video corresponding to a selected song when the user selects a song on the karaoke terminal. This enables dynamic video adjustment according to the user's emotion, providing a richer karaoke experience. It also improves video production efficiency and reduces operational costs.

[0651] "Music data" refers to audio data of music played on a karaoke system.

[0652] "Lyric data" is character information corresponding to a song, and is text data indicating the content of the lyrics.

[0653] An "emotion quantification means" is a method or system that performs a process of analyzing lyric content and assigning a numerical value to a particular emotional category.

[0654] A "generative AI model" is an artificial intelligence model that automatically generates videos from input data such as emotional data.

[0655] "Means for automatically generating videos" refers to the process of using a generative AI model to generate videos based on input data.

[0656] "Means for reviewing videos and fine-tuning them as needed" refers to the processes and tools used to review the content of the videos produced and adjust any inappropriate parts or timing.

[0657] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions, tone of voice, body movements, etc.

[0658] "Means for analyzing in real time" refers to a method for acquiring user emotion data in real time and executing an analysis process.

[0659] The "means for adjusting the content of the video" is a process for appropriately changing the scenes and content of the video being played based on the user's real-time emotional data.

[0660] The "means for uploading to a database" refers to the process or tool for storing the generated final video in a database such as cloud storage or a local server.

[0661] "Means for distributing to user terminals" refers to the process of transferring video data stored in the database to the karaoke terminals used by users.

[0662] The "means for playing video" refers to a function or process for playing video corresponding to a song selected on a user terminal using a display device or audio device.

[0663] The present invention is a system for automatically generating videos to be inserted into a karaoke system based on user emotion recognition. Specific embodiments of the system will be described below.

[0664] System Configuration

[0665] The system mainly consists of the following components:

[0666] 1. Server

[0667] 2. User Device

[0668] 3. Emotion Engine

[0669] Operation overview

[0670] Video and music data collection

[0671] Server: The server collects video and music data from TV dramas and movies from public databases on the Internet and from licensed content providers. This data is prepared as a dataset for training the generative AI model. Specifically, a Python script is used to download the data via an API and save it in storage. This process creates a dataset capable of learning a variety of emotional expressions and scenes.

[0672] Emotional analysis of new karaoke songs

[0673] Server: New karaoke songs and lyrics are input into the system, and natural language processing (NLP) techniques are used to quantify the emotions in the lyrics. This analysis uses emotion recognition libraries (e.g., NLTK, SpaCy). For example, the lyrics of the song "Sakura" are quantified as "70% joy, 30% sadness." This numerical data is used as input data for the generative AI model.

[0674] User Emotion Recognition

[0675] Emotion engine: Analyzes the user's facial expressions, vocal tone, and body movements in real time. This is done using a facial recognition camera, microphone, and motion sensor. The recognized emotional data is sent to a server. Specifically, facial expression analysis is performed using OpenCV, vocal tone analysis using a voice analysis tool, and body movements are analyzed using a motion sensor.

[0676] Optimal video generation

[0677] Server: Real-time emotional data obtained from the emotion engine and quantified emotional data from lyrics are input into the generative AI model. The generative AI model automatically generates optimal video scenes based on this data. For example, for "Sakura," it generates scenes of cherry blossoms in full bloom and a graduation ceremony. The prompt text in this case is as follows:

[0678] Generate an optimal insert video based on the lyrics of "Sakura." The emotional analysis of the lyrics reveals "70% joy, 30% sadness." The emotions recognized from the user's facial expressions, tone of voice, and body movements are "80% joy, 20% sadness." Based on this, create an inspiring video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0679] Video review and fine-tuning

[0680] Server: The server reviews the generated video and manually corrects any inaccuracies or timing adjustments. For example, they use video editing software such as Adobe Premiere Pro to optimize the synchronization between the video and the music.

[0681] Uploading video data

[0682] Server: Uploads the final video to a database and prepares it for distribution to user devices. This data includes meta information (song title, lyrics, sentiment analysis results). The server accesses the database and stores the video data in cloud storage or on a local server.

[0683] Data synchronization to karaoke terminal

[0684] User terminal: Periodically synchronizes with the karaoke system database to receive new video data. The terminal accesses the server and downloads data based on a set schedule.

[0685] Video playback and real-time adjustments

[0686] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. The emotion engine analyzes the user's emotions in real time and adjusts the video content as needed, allowing the user to enjoy a more emotionally immersive karaoke experience.

[0687] Through these steps, dynamic video adjustments based on user emotions are possible, providing a richer karaoke experience, while also improving video production efficiency and reducing operational costs.

[0688] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0689] Step 1:

[0690] Server: Collects video data from TV dramas and movies and music data, and prepares it as a training dataset for the generative AI model.

[0691] Input: Video and music data collected from public databases and content providers on the Internet.

[0692] Data processing: Using a Python script, the database is accessed via API, video and music data is downloaded, and stored in storage.

[0693] Output: The dataset used to train the generative AI model.

[0694] Step 2:

[0695] Server: New karaoke songs and their lyrics are input into the system, and the emotions of the lyrics are analyzed and quantified.

[0696] Input: Karaoke song and its lyrics data.

[0697] Data Arithmetic: Using an emotion recognition library (e.g., NLTK, SpaCy), parse the lyrics and assign numerical values ​​to major emotion categories, e.g., "70% joy, 30% sadness."

[0698] Output: Quantified lyrics emotion data.

[0699] Step 3:

[0700] Emotion engine: Identifies the user's facial expressions, tone of voice, and body movements in real time and sends the emotional data to the server.

[0701] Input: The user's facial expressions, tone of voice, and body movements.

[0702] Data calculation: Facial expressions, voice, and movements are captured using a facial recognition camera, microphone, and motion sensor, and then analyzed using OpenCV and voice analysis tools.

[0703] Output: Real-time sentiment data.

[0704] Step 4:

[0705] Server: Based on the emotional data from the emotion engine and the quantified emotional data from the lyrics, the generative AI model automatically generates optimal video scenes.

[0706] Input: Real-time emotion data from the emotion engine and quantified lyric emotion data.

[0707] Data computation: A generative AI model (e.g., GPT-4, DALL-E) generates video scenes based on prompts.

[0708] Output: The generated video scenes.

[0709] Step 5:

[0710] Server: Review the generated video and manually fine-tune any necessary parts.

[0711] Input: Generated video scenes.

[0712] Data processing: Use video editing software (e.g. Adobe Premiere Pro) to adjust inappropriate parts and timing.

[0713] Output: Final optimized video.

[0714] Step 6:

[0715] Server: Uploads the final video to the database and distributes it to the karaoke terminals.

[0716] Input: The final optimized video and its meta information (song title, lyrics, sentiment analysis results).

[0717] Data processing: Video data is stored in cloud storage or on a local server and prepared for distribution.

[0718] Output: Video data uploaded to the database.

[0719] Step 7:

[0720] Terminal: Synchronizes with the karaoke system database and receives new video data.

[0721] Input: New video data from the database.

[0722] Data processing: Through a periodic synchronization process, the server is accessed, video data is downloaded, and stored in the internal storage.

[0723] Output: New video data saved on your device.

[0724] Step 8:

[0725] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. An emotion recognition engine adjusts the video content in real time as needed.

[0726] Input: User-selected song, real-time emotional data.

[0727] Data calculation: During video playback, the emotion recognition engine analyzes the user's emotions and adjusts the video content and timing accordingly.

[0728] Output: Real-time adjusted video playback provides a more emotionally immersive karaoke experience.

[0729] (Application example 2)

[0730] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] Modern advertising videos are unable to adapt their content in real time to the viewer's emotions, making it difficult to provide an optimal visual and emotional experience for each individual user, which limits the effectiveness of advertising and reduces viewer engagement.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0733] In this invention, the server includes: means for inputting music data and lyric data and quantifying the emotion of the lyrics; means for automatically generating videos using a generative AI model using the quantified emotion data as input data; means for reviewing the generated videos and fine-tuning them as necessary; means for uploading the final videos to a database and distributing them to user terminals; means for acquiring the user's facial expressions, tone of voice, and movements in real time and analyzing their emotions; and means for the generative AI model to dynamically change the video content in real time based on the emotion data. This makes it possible to generate advertising videos that adapt to the user's emotions in real time, improve viewer engagement, and maximize advertising effectiveness.

[0734] "Music data" refers to the audio information of a musical work and its associated metadata.

[0735] "Lyric data" refers to text information about lyrics corresponding to a song.

[0736] The "emotion engine" is a function that analyzes emotions in real time from the user's facial expressions, voice, movements, etc.

[0737] "Numericalized emotion data" refers to data that expresses emotions as numerical values.

[0738] A "generative AI model" is an artificial intelligence model that generates new data or content based on input data.

[0739] "Automatic animation generation means" refers to a method or device for automatically generating animation using input data.

[0740] A "database" is a system for efficiently storing, managing, and retrieving data.

[0741] A "user terminal" is an electronic device that is directly operated by a user.

[0742] "Tweaking tool" refers to a method or device for manually adjusting generated content.

[0743] "Facial expression recognition" is a technology that analyzes a user's facial expressions and identifies their emotions based on them.

[0744] "Tone of voice" refers to characteristics such as pitch, intensity, and speed of voice, and is an element that conveys the speaker's emotions and intentions.

[0745] "Real-time acquisition" refers to the acquisition and processing of data immediately, without delay.

[0746] "Dynamic change" means that a system or content changes automatically over time or in response to events.

[0747] An "advertising video" is video content created to promote a particular product or service.

[0748] "Viewer's emotions" refer to the emotions and feelings felt by a user watching video content.

[0749] The present invention relates to a system for dynamically generating and adjusting advertising videos in real time based on user emotions. Specific embodiments of the system are described below.

[0750] System configuration

[0751] The system mainly consists of the following components:

[0752] 1. Server

[0753] 2. User Device

[0754] 3. Emotion Engine

[0755] Program processing overview

[0756] Hardware and Software Configuration

[0757] Hardware:

[0758] Smartphone camera

[0759] Smartphone microphone

[0760] software:

[0761] EmotionEngine (emotion analysis engine)

[0762] Video Generator (generative AI model)

[0763] Database management systems (e.g., SQLite)

[0764] System Operation

[0765] 1. The server first inputs the music data and lyrics data and quantifies the emotion of the lyrics.

[0766] 2. Using the quantified emotional data as input data, a generative AI model is used to automatically generate videos.

[0767] 3. Review the generated video on the server side and make any necessary adjustments.

[0768] 4. Upload the final video to the database and distribute it to the user's device.

[0769] 5. While the user is watching the advertising video, the user's device will use a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements.

[0770] 6. The emotion engine analyzes the user's emotions based on the acquired data and generates quantified emotion data.

[0771] 7. The generative AI model dynamically changes video content based on emotional data captured in real time.

[0772] 8. The user device plays the adjusted video in real time.

[0773] Specific examples

[0774] For example, consider an advertising video for a "new sports car." This advertisement dynamically changes to emphasize the speed of the car when the user is excited, and to show scenes of the car traveling through beautiful scenery when the user is relaxed. Specific examples of prompt sentences are as follows:

[0775] Example prompt:

[0776] "For an ad for a new sports car, generate a video that shows a car speeding along if the user is excited, or a video that shows a car driving through beautiful scenery if the user is relaxed."

[0777] This system can adapt to user emotions in real time and improve viewer engagement. In particular, in the case of advertising videos, dynamically changing content to match user emotions can maximize advertising effectiveness.

[0778] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0779] Step 1:

[0780] The server inputs music data and lyrics data. The server then quantifies the emotions in the lyrics and generates emotional data. Specifically, it uses natural language processing to analyze the main emotional elements of the lyrics and quantifies them, such as "70% joy, 30% sadness."

[0781] Step 2:

[0782] The server uses the quantified emotional data as input and automatically generates videos using a generative AI model. Based on the input emotional data, it constructs videos from highly relevant scenes. For example, if the emotional data is "70% joy, 30% sadness," it selects appropriate scenes (such as cherry blossoms in full bloom or a graduation ceremony) and generates a video.

[0783] Step 3:

[0784] The server reviews the generated video and makes any necessary adjustments, such as manually correcting scene timing or inappropriate parts of the generated video. As a result of the review, the final video file is generated.

[0785] Step 4:

[0786] The server uploads the final video to a database and distributes it to the user's device. Specifically, the video data is stored in the database along with meta information (song title, lyrics, and sentiment analysis results) and made accessible from the client device.

[0787] Step 5:

[0788] While the user is watching the advertising video, the device uses a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements. Specifically, the device's camera captures video data and the microphone captures audio data.

[0789] Step 6:

[0790] Based on the data acquired by the emotion engine, the user's emotions are analyzed and quantified emotional data is generated. Specifically, emotional data such as "80% excited, 20% relaxed" is generated from the user's facial expressions and tone of voice.

[0791] Step 7:

[0792] The generative AI model dynamically changes the video content based on emotional data acquired in real time. For example, if the user is excited, it will switch to scenes that emphasize the speed of cars, and if the user is relaxed, it will switch to scenes with beautiful scenery.

[0793] Step 8:

[0794] The device then plays the adjusted video in real time, ultimately displaying advertising videos that match the user's emotions, improving viewer engagement and maximizing advertising effectiveness.

[0795] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0796] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0797] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0798] [Third embodiment]

[0799] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0800] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0801] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0802] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0803] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0804] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0805] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0806] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0807] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0808] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0809] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0810] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0811] The present invention provides a system for automatically generating karaoke insert videos, and a specific embodiment thereof will be described below.

[0812] System Overview

[0813] The system is mainly composed of three components: a server, a device, and a user. The server automatically generates videos using a generative AI model and manages distribution to the device. The device provides the karaoke execution environment and an interface for users to enjoy karaoke. Users select a song and sing karaoke while visually enjoying the inserted video that matches the song.

[0814] Program processing overview

[0815] 1. The server collects video data from TV dramas and movies and corresponding music data, and prepares it as a training dataset for the generative AI model, which then learns the relationship between video scenes and music.

[0816] 2. The server inputs the new karaoke song and its lyrics, analyzes the emotions in the lyrics, and converts them into numerical values. For example, the main emotions in the lyrics are quantified as "joy 60%, sadness 40%."

[0817] 3. The server inputs the quantified emotional data into a generative AI model and automatically generates a video based on the emotional parameters. For example, for a song with a spring cherry blossom theme, a video including scenes of cherry blossoms in full bloom and students graduating will be generated.

[0818] 4. The server reviews the generated video and manually adjusts it as needed, for example by correcting the timing of certain scenes to optimize the synchronization between the video and the music.

[0819] 5. The server uploads the final video to a database and delivers it to the karaoke terminals, with the video data being transferred over a high-speed, stable connection.

[0820] 6. When the user selects a song, the device plays a video corresponding to that song. For example, if the user selects the song "Sakura," a pre-generated video of cherry blossoms in full bloom will be played.

[0821] 7. Users can visually enjoy the inserted video while singing the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0822] Specific examples

[0823] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0824] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0825] 2. Server: The quantified emotional data is input into the generative AI model, and a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony that match the content of the lyrics.

[0826] 3. Server: Preview the generated video and manually fine-tune the timing of scenes, further refining the synchronization between the video and music.

[0827] 4. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0828] 5. Device: When the user selects "Sakura," the video corresponding to the selected song will be played.

[0829] 6. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of graduation ceremonies.

[0830] In this way, this system can provide insert videos that match the worldview of the song when users enjoy karaoke, improving the visual entertainment value.In addition, by using a generative AI model, it is possible to significantly reduce the cost of video production and the burden of copyright fees.

[0831] The processing flow will be explained below.

[0832] Step 1:

[0833] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[0834] Step 2:

[0835] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and irrelevant data.

[0836] Step 3:

[0837] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and letting it learn the associations between videos and music based on the dataset.

[0838] Step 4:

[0839] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[0840] Step 5:

[0841] The server analyzes the emotions in the lyrics and converts them into numerical values, such as joy, anger, sadness, and pleasure. In this case, the result might be "70% joy, 30% sadness."

[0842] Step 6:

[0843] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[0844] Step 7:

[0845] The server reviews the generated video and manually adjusts it as needed, correcting scene timing or other inappropriate parts.

[0846] Step 8:

[0847] The server uploads the final video to a database, including meta information about the video data (such as song title, lyrics, and sentiment analysis results).

[0848] Step 9:

[0849] The terminal synchronizes with the karaoke system's database to receive new video data, ensuring that data transfer is fast and stable.

[0850] Step 10:

[0851] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[0852] Step 11:

[0853] The device plays the video corresponding to the song selected by the user. While the song "Sakura" is playing, scenes of cherry blossoms in full bloom and a graduation ceremony are displayed.

[0854] Step 12:

[0855] Users can visually enjoy the inserted video while singing along to the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[0856] Example 1

[0857] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0858] In conventional karaoke systems, manually creating videos that match a song is time-consuming and costly, making it difficult to provide users with a satisfying entertainment experience. Furthermore, it is not possible to generate videos that match the emotions of the lyrics in real time, making it difficult to effectively convey the worldview of the song. To solve these issues, there is a need for a system that can efficiently and automatically generate videos that match a song and provide users with a high-quality karaoke experience.

[0859] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0860] In this invention, the server includes means for inputting music data and text data and quantifying emotions in the text data, means for using the quantified emotional data as input data and automatically generating videos using a generative AI model, means for reviewing the generated videos and fine-tuning them as necessary, means for uploading the final videos to a database and distributing them to a user terminal, and means for playing videos corresponding to the selected song when the user selects a song on the karaoke terminal. This makes it possible to efficiently and automatically generate videos suited to songs and provide users with a high-quality karaoke experience.

[0861] "Music data" is digital data of audio information that makes up music.

[0862] "Character data" is digital data that includes text information such as lyrics.

[0863] A "means for quantifying emotions" is a method or device that analyzes the content of character data and expresses its emotional elements as numerical values.

[0864] A "generative AI model" is an algorithm or program that uses machine learning and deep learning to automatically generate images.

[0865] "Means for automatically generating video" refers to a method or device that uses quantified emotional data and utilizes a generative AI model to create video.

[0866] The "reviewing means" refers to a method or device for checking the generated video and correcting it if necessary.

[0867] A "database" is an information storage system for storing and managing video and music data.

[0868] A "user terminal" is a device or apparatus that a user uses to enjoy karaoke.

[0869] A "karaoke terminal" is a device or apparatus dedicated to playing karaoke songs and displaying images.

[0870] The "video playback means" refers to a method or device for displaying video in accordance with the song selected on the karaoke terminal.

[0871] The present invention is a system for automatically generating images suitable for karaoke songs. A specific embodiment of this system will be described below.

[0872] The system mainly consists of three components: a server, a device, and a user. The server has a high-performance processor (e.g., an NVIDIA GPU) and large storage capacity, and uses machine learning frameworks such as TensorFlow or PyTorch to train and run the generative AI model. The device is an Android or iOS device, or a dedicated karaoke machine with built-in video playback software. The user uses a smartphone or a device in a karaoke booth.

[0873] The server first collects video data from TV dramas and movies and the corresponding music data. This creates a dataset for learning the association between video scenes and music. Specifically, the server obtains video and music data from the Internet or a dedicated database, and then formats them.

[0874] Next, the server inputs the new song and its text data. The text data includes information such as lyrics and title. This is analyzed using natural language processing (NLP) tools to quantify the emotions. For example, the lyrics of the song "Sakura" are quantified as "60% joy, 40% sadness." The analysis results are input into the generative AI model as prompt sentences.

[0875] Here is an example prompt: "Please analyze the emotions from the lyrics data of the specified song. Based on the analysis results, please generate an image that matches the lyrics. The emotional data is '70% joy, 30% sadness'."

[0876] The server uses the emotion data as input and automatically generates video using a generative AI model, which then creates a video based on the emotion data, generating scenes such as cherry blossoms in full bloom or a graduation ceremony.

[0877] The generated video is previewed on the server, and manual adjustments are made as needed, such as correcting the timing of specific scenes to optimize the synchronization between the video and music. Once the final adjustments are complete, the video is uploaded to the database.

[0878] The server delivers the videos uploaded to the database to the user's device. This delivery is done via a high-speed, stable connection. When the user selects a song, the device plays the video corresponding to that song in real time. For example, if the user selects "Sakura," a video of cherry blossoms in full bloom will be played as the intro begins.

[0879] Users can visually enjoy the video while singing the song. Because the video scenes are linked to the lyrics, users can immerse themselves deeply in the world of the song while singing. For example, while singing the song "Sakura," users can visually enjoy scenes of cherry blossoms in full bloom in spring or a graduation ceremony. In this way, users can be provided with a high-quality karaoke experience.

[0880] This system significantly reduces the cost of video production while providing users with video that matches the worldview of the song.

[0881] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0882] Step 1:

[0883] The server collects video data of TV dramas and movies and corresponding music data from the Internet or a dedicated database. This data is used as a training dataset for the generative AI model. The input data is the video and music data of TV dramas and movies, and the output data is a formatted training dataset. Specifically, the server obtains a scene from a movie and its background music and formats it into a dataset.

[0884] Step 2:

[0885] The server receives the newly added song and its lyrics data. It then uses a natural language processing (NLP) tool to quantify the emotion of the lyrics. The input data is the song and lyrics data, and the output data is the quantified emotion data. Specifically, the server analyzes the lyrics of the song "Sakura" and generates emotion data of "60% joy, 40% sadness."

[0886] Step 3:

[0887] The server converts the quantified emotional data into a prompt sentence and inputs it into the generative AI model. At this time, the generative AI model automatically generates a video based on the emotional parameters. The input data is the quantified emotional data, and the output data is the generated video. Specifically, the generated video includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0888] Step 4:

[0889] The server reviews the generated footage and manually adjusts it as necessary. The input data is the generated footage, and the output data is the adjusted footage. Specifically, the server corrects the timing of specific scenes and optimizes the synchronization between the footage and the music.

[0890] Step 5:

[0891] The server uploads the final video to a database and distributes it to the user's device. The input data is the adjusted video, and the output data is the video stored in the database and the distributed video data. Specifically, the video data is stored in cloud storage and a link to it is sent to the device.

[0892] Step 6:

[0893] When the user selects a song, the device plays the video delivered from the server in real time. The input data is the user's song selection information and the delivered video data, and the output data is the video to be played. Specifically, when the user selects "Sakura," the intro starts and a video of cherry blossoms in full bloom is played at the same time.

[0894] Step 7:

[0895] Users visually enjoy the video while singing the song. The input data are karaoke songs and video data, and the output data is the user's karaoke experience. Specifically, users can sing "Sakura" while watching a touching graduation scene, enriching their karaoke experience.

[0896] (Application example 1)

[0897] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0898] There is a demand for combining music and video content to provide new content experiences that users can enjoy even more. However, creating visuals that match music is costly and requires a lot of effort. In addition, manual video production is time-consuming, and it is difficult to optimize the synchronization between music and video during the process. This issue needs to be resolved.

[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0900] In this invention, the server includes means for inputting music data and lyric data and quantifying the emotion of the lyrics, means for automatically generating videos using a generative AI model using the quantified emotion data as input data, means for reviewing the generated videos and fine-tuning them as necessary, and means for playing videos corresponding to the selected songs when a user selects a song from a content distribution service. This improves the user experience, significantly reduces the cost and effort of video production, and enables highly accurate synchronization of music and video.

[0901] "Music data" refers to music information stored in digital format.

[0902] "Lyrics data" refers to the text information of the lyrics of a song stored in digital format.

[0903] "Emotional data" refers to the quantified information of emotional elements analyzed from lyrics.

[0904] A "generative AI model" refers to a machine learning model that uses artificial intelligence to generate new outputs from input data.

[0905] "Means for automatically generating videos" refers to the process of using an AI model to automatically create videos based on emotional data.

[0906] "Means of reviewing and fine-tuning as necessary" refers to the process of checking the generated video and manually correcting it as necessary.

[0907] "Content distribution service" refers to a service that provides digital content to users via the Internet.

[0908] "User terminal" refers to a device that a user uses to connect to the Internet, such as a smartphone or tablet.

[0909] This invention will explain a method for constructing a system in which music data and lyrics data are input, and related inserted animation is automatically generated and provided to the user.

[0910] System configuration

[0911] The system consists of the following hardware and software:

[0912] Server: Cloud server. Manages music data, analyzes lyrics sentiment, runs generative AI models, and manages generated videos.

[0913] Generative AI models: Use generative models such as GPT-4 and DALL-E.

[0914] Lyric analysis software: A text analysis tool that uses natural language processing technology.

[0915] Streaming software: Use HLS (HTTP Live Streaming) or similar.

[0916] User device: The device through which the user receives content, such as a smartphone or tablet.

[0917] Operational Overview

[0918] 1. The user selects a song

[0919] The user selects a song using the application of the content distribution service.

[0920] 2. The server retrieves the music data and lyrics data

[0921] Based on the song ID selected by the user, the server retrieves the song data and lyrics data.

[0922] 3. Lyrics Analysis

[0923] The server inputs the acquired lyrics data into lyrics analysis software and quantifies the emotions of the lyrics. For example, the main emotional elements of the lyrics are quantified as "joy 60% and sadness 40%."

[0924] 4. Visual Generation

[0925] The server inputs the quantified emotional data into a generative AI model to generate video visuals based on the emotions. For example, based on emotional data of "60% joy, 40% sadness," a video containing scenes of cherry blossoms in full bloom and students graduating might be generated.

[0926] 5. Review and fine-tune

[0927] The video generated on the server is previewed and manually tweaked as needed, correcting the timing of specific scenes and optimizing the synchronization of the video and music.

[0928] 6. Video distribution

[0929] The final fine-tuned video is delivered to the user's device using streaming software (such as HLS).

[0930] Specific examples

[0931] The prompt text and system behavior when the user selects the song "Spring Breeze" in the app are shown below.

[0932] Prompt statement

[0933] Visualize a spring landscape with emotional parameters of 70% hope and 30% joy.

[0934] System operation example

[0935] 1. When the user selects the song "Spring Breeze," the server retrieves the corresponding song data and lyrics data.

[0936] 2. Analyze the acquired lyrics data and quantify the emotional data (hope 70%, joy 30%).

[0937] 3. This emotional data is input into a generative AI model to generate images that express hope and joy, such as the fresh greenery of spring and bright skies.

[0938] 4. The generated video is reviewed on the server side, and the timing is fine-tuned to achieve optimal synchronization.

[0939] 5. The final video will be delivered to the user’s smartphone, where they can enjoy listening to the song “Spring Breeze.”

[0940] In this way, the synchronization between the music and the video can be improved, providing the user with a more attractive content experience.

[0941] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0942] Step 1:

[0943] The server receives an operation by the user to select a song and acquires the song ID.

[0944] Input: User-selected song ID

[0945] Output: Song ID is passed to the system

[0946] How it works: A user selects the song "Spring Breeze" on the app. The server detects the selection and retrieves the song ID (e.g., "12345").

[0947] Step 2:

[0948] The server uses the song ID to obtain the song data and lyrics data.

[0949] Input: Song ID

[0950] Output: Music data and lyrics data

[0951] Operation: The server accesses the database and retrieves the music data and lyrics data corresponding to the music ID "12345." At this time, the music data is retrieved in music file format, and the lyrics data is retrieved in text format.

[0952] Step 3:

[0953] The server analyzes the acquired lyrics data and quantifies the emotion of the lyrics.

[0954] Input: Lyrics data

[0955] Output: Emotion data (e.g., hope 70%, joy 30%)

[0956] How it works: The server inputs the lyrics data into natural language processing software, which analyzes the lyrics text. This analysis quantifies the emotional elements contained in the lyrics and generates emotional data, such as "hope 70%, joy 30%."

[0957] Step 4:

[0958] The server inputs the quantified emotional data into a generative AI model and automatically generates videos based on the emotions.

[0959] Input: Emotion data

[0960] Output: Generated video

[0961] How it works: The server inputs the prompt "Please visualize a spring landscape with emotional parameters of 70% hope and 30% joy" into a generative AI model (e.g., DALL-E or GPT-4), and generates a video based on this. For example, a visual containing fresh spring greenery and a bright sky is generated.

[0962] Step 5:

[0963] The server reviews the generated video and makes any necessary adjustments.

[0964] Input: Generated video

[0965] Output: Final fine-tuned video

[0966] How it works: Preview the video generated on the server, check the timing and sequence of scenes, and if necessary, manually tweak specific scenes and timing to optimize the synchronization of music and visuals.

[0967] Step 6:

[0968] The server then delivers the final, finely tuned video to the user's device via streaming software.

[0969] Input: Final tweaked video

[0970] Output: Streaming video to user devices

[0971] How it works: The server uses HLS streaming software to encode the final video in real time and delivers it to the user's smartphone, tablet, or other device. Users can then watch the video corresponding to the song they selected on the app.

[0972] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0973] The present invention provides a system for automatically generating karaoke insert videos, which incorporates a technology that combines an emotion engine that recognizes the user's emotions. Specific embodiments of the technology are described below.

[0974] System configuration

[0975] The system mainly consists of the following components:

[0976] 1. Server

[0977] 2. User Device

[0978] 3. Emotion Engine

[0979] Program processing overview

[0980] 1. Server: Collects video and music data from TV dramas and movies and prepares it as a training dataset for the generative AI model, allowing the model to learn the relationship between video scenes and music.

[0981] 2. Server: A new karaoke song and its lyrics are input into the system, and the emotions in the lyrics are analyzed and quantified. For example, the main emotions in the lyrics are quantified as "joy 60% and sadness 40%."

[0982] 3. Server: The emotion engine recognizes the user's emotions in real time and feeds that emotion data back to the generative AI model. The emotion engine recognizes emotions from the user's facial expressions, tone of voice, body movements, etc.

[0983] 4. Server: Based on the emotional data obtained from the emotion engine and the quantified emotional data of the lyrics, the generative AI model automatically generates the optimal video. For example, if the user's emotions are "70% joy, 30% sadness," it will generate a video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0984] 5. Server: Review the generated video and manually tweak it as needed, such as correcting scene timing or inappropriate parts.

[0985] 6. Server: Uploads the final video to the database and distributes it to the karaoke terminals, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[0986] 7. Terminal: Synchronizes with the karaoke system database and receives new video data, ensuring fast and stable data transfer.

[0987] 8. User: When a song is selected from the karaoke song list, a video corresponding to the song will be played, and the video will be adjusted in real time based on the user's emotions.

[0988] 9. Device: The device plays a video corresponding to the song selected by the user, and the emotion engine analyzes emotions in real time and adjusts the content of the video as needed.

[0989] 10. Users: They can visually enjoy the inserted video while singing the song. By adjusting the video based on the user's emotions, they can more easily immerse themselves in the world of the song, improving the karaoke experience.

[0990] Specific examples

[0991] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[0992] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[0993] 2. Emotion Engine: While the user is singing "Sakura," the system recognizes the user's emotions from their facial expressions, tone of voice, and body movements. For example, if the user is singing happily, it will be recognized as "80% joy, 20% sadness."

[0994] 3. Server: The emotion data recognized by the emotion engine is fed back to the AI ​​model, and combined with the emotional data from the lyrics, a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[0995] 4. Server: Preview the generated video and manually fine-tune the timing of scenes to optimize the synchronization between the video and music.

[0996] 5. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[0997] 6. Device: When the user selects "Sakura," a video adjusted in real time based on the user's emotions is played.

[0998] 7. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of a graduation ceremony. In addition, the content of the video is adjusted in real time to match the user's emotions, allowing for a more emotionally immersive experience.

[0999] In this way, the system recognizes the user's emotions and uses that emotional data to adjust the video in real time, providing a visually enjoyable karaoke experience. Furthermore, the use of generative AI models can significantly reduce video production costs and copyright fees.

[1000] The processing flow will be explained below.

[1001] Step 1:

[1002] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[1003] Step 2:

[1004] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and inappropriate data.

[1005] Step 3:

[1006] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and training the model to recognize associations between videos and music based on the dataset.

[1007] Step 4:

[1008] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[1009] Step 5:

[1010] The server analyzes the emotions in the lyrics and quantifies emotions such as joy, anger, sadness, and happiness. In this case, the result might be "70% joy, 30% sadness."

[1011] Step 6:

[1012] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[1013] Step 7:

[1014] The server reviews the generated video and manually adjusts it to correct timing or inappropriate parts of the scene.

[1015] Step 8:

[1016] The server uploads the final video to a database, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[1017] Step 9:

[1018] The terminal synchronizes with the karaoke system's database and receives new video data, ensuring fast and stable data transfer.

[1019] Step 10:

[1020] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[1021] Step 11:

[1022] The device plays a video corresponding to the song selected by the user. As the "Sakura" video plays, the emotion engine analyzes the user's emotions in real time.

[1023] Step 12:

[1024] The emotion engine recognizes the user's emotions in real time from their facial expressions, tone of voice, and body movements.

[1025] Step 13:

[1026] The server then feeds the emotion data recognized by the emotion engine back to the generative AI model, adjusting the video content in real time. For example, if the user is having fun, it adds a cheerful scene.

[1027] Step 14:

[1028] The device displays real-time updated video to the user, with the video content tailored to the user's current emotions.

[1029] Step 15:

[1030] Users can enjoy visually tailored videos while singing along, which helps users immerse themselves in the world of the song and enhances the karaoke experience.

[1031] In this way, by incorporating an emotion engine, it becomes possible to adjust the video content according to the user's real-time emotions, providing a more immersive karaoke experience.

[1032] Example 2

[1033] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1034] In conventional karaoke systems, the inserted video is fixed, making it impossible to dynamically adjust the video according to the user's emotions or singing style, resulting in a limited karaoke experience. In addition, creating fixed videos requires a great deal of effort and cost, making efficient operation difficult.

[1035] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1036] In this invention, the server includes: a means for inputting music data and lyric data and quantifying the emotion of the lyrics; a means for automatically generating videos using a generative AI model using the quantified emotion data as input data; a means for reviewing the generated videos and fine-tuning them as necessary; an emotion recognition means for analyzing the user's facial expressions, tone of voice, and body movements in real time; a means for feeding back the analyzed emotion data in real time to the generative AI model and adjusting the content of the videos; a means for uploading the final videos to a database and distributing them to a user terminal; and a means for playing a video corresponding to a selected song when the user selects a song on the karaoke terminal. This enables dynamic video adjustment according to the user's emotion, providing a richer karaoke experience. It also improves video production efficiency and reduces operational costs.

[1037] "Music data" refers to audio data of music played on a karaoke system.

[1038] "Lyric data" is character information corresponding to a song, and is text data indicating the content of the lyrics.

[1039] An "emotion quantification means" is a method or system that performs a process of analyzing lyric content and assigning a numerical value to a particular emotional category.

[1040] A "generative AI model" is an artificial intelligence model that automatically generates videos from input data such as emotional data.

[1041] "Means for automatically generating videos" refers to the process of using a generative AI model to generate videos based on input data.

[1042] "Means for reviewing videos and fine-tuning them as needed" refers to the processes and tools used to review the content of the videos produced and adjust any inappropriate parts or timing.

[1043] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions, tone of voice, body movements, etc.

[1044] "Means for analyzing in real time" refers to a method for acquiring user emotion data in real time and executing an analysis process.

[1045] The "means for adjusting the content of the video" is a process for appropriately changing the scenes and content of the video being played based on the user's real-time emotional data.

[1046] The "means for uploading to a database" refers to the process or tool for storing the generated final video in a database such as cloud storage or a local server.

[1047] "Means for distributing to user terminals" refers to the process of transferring video data stored in the database to the karaoke terminals used by users.

[1048] The "means for playing video" refers to a function or process for playing video corresponding to a song selected on a user terminal using a display device or audio device.

[1049] The present invention is a system for automatically generating videos to be inserted into a karaoke system based on user emotion recognition. Specific embodiments of the system will be described below.

[1050] System Configuration

[1051] The system mainly consists of the following components:

[1052] 1. Server

[1053] 2. User Device

[1054] 3. Emotion Engine

[1055] Operation overview

[1056] Video and music data collection

[1057] Server: The server collects video and music data from TV dramas and movies from public databases on the Internet and from licensed content providers. This data is prepared as a dataset for training the generative AI model. Specifically, a Python script is used to download the data via an API and save it in storage. This process creates a dataset capable of learning a variety of emotional expressions and scenes.

[1058] Emotional analysis of new karaoke songs

[1059] Server: New karaoke songs and lyrics are input into the system, and natural language processing (NLP) techniques are used to quantify the emotions in the lyrics. This analysis uses emotion recognition libraries (e.g., NLTK, SpaCy). For example, the lyrics of the song "Sakura" are quantified as "70% joy, 30% sadness." This numerical data is used as input data for the generative AI model.

[1060] User Emotion Recognition

[1061] Emotion engine: Analyzes the user's facial expressions, vocal tone, and body movements in real time. This is done using a facial recognition camera, microphone, and motion sensor. The recognized emotional data is sent to a server. Specifically, facial expression analysis is performed using OpenCV, vocal tone analysis using a voice analysis tool, and body movements are analyzed using a motion sensor.

[1062] Optimal video generation

[1063] Server: Real-time emotional data obtained from the emotion engine and quantified emotional data from lyrics are input into the generative AI model. The generative AI model automatically generates optimal video scenes based on this data. For example, for "Sakura," it generates scenes of cherry blossoms in full bloom and a graduation ceremony. The prompt text in this case is as follows:

[1064] Generate an optimal insert video based on the lyrics of "Sakura." The emotional analysis of the lyrics reveals "70% joy, 30% sadness." The emotions recognized from the user's facial expressions, tone of voice, and body movements are "80% joy, 20% sadness." Based on this, create an inspiring video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[1065] Video review and fine-tuning

[1066] Server: The server reviews the generated video and manually corrects any inaccuracies or timing adjustments. For example, they use video editing software such as Adobe Premiere Pro to optimize the synchronization between the video and the music.

[1067] Uploading video data

[1068] Server: Uploads the final video to a database and prepares it for distribution to user devices. This data includes meta information (song title, lyrics, sentiment analysis results). The server accesses the database and stores the video data in cloud storage or on a local server.

[1069] Data synchronization to karaoke terminal

[1070] User terminal: Periodically synchronizes with the karaoke system database to receive new video data. The terminal accesses the server and downloads data based on a set schedule.

[1071] Video playback and real-time adjustments

[1072] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. The emotion engine analyzes the user's emotions in real time and adjusts the video content as needed, allowing the user to enjoy a more emotionally immersive karaoke experience.

[1073] Through these steps, dynamic video adjustments based on user emotions are possible, providing a richer karaoke experience, while also improving video production efficiency and reducing operational costs.

[1074] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1075] Step 1:

[1076] Server: Collects video data from TV dramas and movies and music data, and prepares it as a training dataset for the generative AI model.

[1077] Input: Video and music data collected from public databases and content providers on the Internet.

[1078] Data processing: Using a Python script, the database is accessed via API, video and music data is downloaded, and stored in storage.

[1079] Output: The dataset used to train the generative AI model.

[1080] Step 2:

[1081] Server: New karaoke songs and their lyrics are input into the system, and the emotions of the lyrics are analyzed and quantified.

[1082] Input: Karaoke song and its lyrics data.

[1083] Data Arithmetic: Using an emotion recognition library (e.g., NLTK, SpaCy), parse the lyrics and assign numerical values ​​to major emotion categories, e.g., "70% joy, 30% sadness."

[1084] Output: Quantified lyrics emotion data.

[1085] Step 3:

[1086] Emotion engine: Identifies the user's facial expressions, tone of voice, and body movements in real time and sends the emotional data to the server.

[1087] Input: The user's facial expressions, tone of voice, and body movements.

[1088] Data calculation: Facial expressions, voice, and movements are captured using a facial recognition camera, microphone, and motion sensor, and then analyzed using OpenCV and voice analysis tools.

[1089] Output: Real-time sentiment data.

[1090] Step 4:

[1091] Server: Based on the emotional data from the emotion engine and the quantified emotional data from the lyrics, the generative AI model automatically generates optimal video scenes.

[1092] Input: Real-time emotion data from the emotion engine and quantified lyric emotion data.

[1093] Data computation: A generative AI model (e.g., GPT-4, DALL-E) generates video scenes based on prompts.

[1094] Output: The generated video scenes.

[1095] Step 5:

[1096] Server: Review the generated video and manually fine-tune any necessary parts.

[1097] Input: Generated video scenes.

[1098] Data processing: Use video editing software (e.g. Adobe Premiere Pro) to adjust inappropriate parts and timing.

[1099] Output: Final optimized video.

[1100] Step 6:

[1101] Server: Uploads the final video to the database and distributes it to the karaoke terminals.

[1102] Input: The final optimized video and its meta information (song title, lyrics, sentiment analysis results).

[1103] Data processing: Video data is stored in cloud storage or on a local server and prepared for distribution.

[1104] Output: Video data uploaded to the database.

[1105] Step 7:

[1106] Terminal: Synchronizes with the karaoke system database and receives new video data.

[1107] Input: New video data from the database.

[1108] Data processing: Through a periodic synchronization process, the server is accessed, video data is downloaded, and stored in the internal storage.

[1109] Output: New video data saved on your device.

[1110] Step 8:

[1111] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. An emotion recognition engine adjusts the video content in real time as needed.

[1112] Input: User-selected song, real-time emotional data.

[1113] Data calculation: During video playback, the emotion recognition engine analyzes the user's emotions and adjusts the video content and timing accordingly.

[1114] Output: Real-time adjusted video playback provides a more emotionally immersive karaoke experience.

[1115] (Application example 2)

[1116] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1117] Modern advertising videos are unable to adapt their content in real time to the viewer's emotions, making it difficult to provide an optimal visual and emotional experience for each individual user, which limits the effectiveness of advertising and reduces viewer engagement.

[1118] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1119] In this invention, the server includes: means for inputting music data and lyric data and quantifying the emotion of the lyrics; means for automatically generating videos using a generative AI model using the quantified emotion data as input data; means for reviewing the generated videos and fine-tuning them as necessary; means for uploading the final videos to a database and distributing them to user terminals; means for acquiring the user's facial expressions, tone of voice, and movements in real time and analyzing their emotions; and means for the generative AI model to dynamically change the video content in real time based on the emotion data. This makes it possible to generate advertising videos that adapt to the user's emotions in real time, improve viewer engagement, and maximize advertising effectiveness.

[1120] "Music data" refers to the audio information of a musical work and its associated metadata.

[1121] "Lyric data" refers to text information about lyrics corresponding to a song.

[1122] The "emotion engine" is a function that analyzes emotions in real time from the user's facial expressions, voice, movements, etc.

[1123] "Numericalized emotion data" refers to data that expresses emotions as numerical values.

[1124] A "generative AI model" is an artificial intelligence model that generates new data or content based on input data.

[1125] "Automatic animation generation means" refers to a method or device for automatically generating animation using input data.

[1126] A "database" is a system for efficiently storing, managing, and retrieving data.

[1127] A "user terminal" is an electronic device that is directly operated by a user.

[1128] "Tweaking tool" refers to a method or device for manually adjusting generated content.

[1129] "Facial expression recognition" is a technology that analyzes a user's facial expressions and identifies their emotions based on them.

[1130] "Tone of voice" refers to characteristics such as pitch, intensity, and speed of voice, and is an element that conveys the speaker's emotions and intentions.

[1131] "Real-time acquisition" refers to the acquisition and processing of data immediately, without delay.

[1132] "Dynamic change" means that a system or content changes automatically over time or in response to events.

[1133] An "advertising video" is video content created to promote a particular product or service.

[1134] "Viewer's emotions" refer to the emotions and feelings felt by a user watching video content.

[1135] The present invention relates to a system for dynamically generating and adjusting advertising videos in real time based on user emotions. Specific embodiments of the system are described below.

[1136] System configuration

[1137] The system mainly consists of the following components:

[1138] 1. Server

[1139] 2. User Device

[1140] 3. Emotion Engine

[1141] Program processing overview

[1142] Hardware and Software Configuration

[1143] Hardware:

[1144] Smartphone camera

[1145] Smartphone microphone

[1146] software:

[1147] EmotionEngine (emotion analysis engine)

[1148] Video Generator (generative AI model)

[1149] Database management systems (e.g., SQLite)

[1150] System Operation

[1151] 1. The server first inputs the music data and lyrics data and quantifies the emotion of the lyrics.

[1152] 2. Using the quantified emotional data as input data, a generative AI model is used to automatically generate videos.

[1153] 3. Review the generated video on the server side and make any necessary adjustments.

[1154] 4. Upload the final video to the database and distribute it to the user's device.

[1155] 5. While the user is watching the advertising video, the user's device will use a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements.

[1156] 6. The emotion engine analyzes the user's emotions based on the acquired data and generates quantified emotion data.

[1157] 7. The generative AI model dynamically changes video content based on emotional data captured in real time.

[1158] 8. The user device plays the adjusted video in real time.

[1159] Specific examples

[1160] For example, consider an advertising video for a "new sports car." This advertisement dynamically changes to emphasize the speed of the car when the user is excited, and to show scenes of the car traveling through beautiful scenery when the user is relaxed. Specific examples of prompt sentences are as follows:

[1161] Example prompt:

[1162] "For an ad for a new sports car, generate a video that shows a car speeding along if the user is excited, or a video that shows a car driving through beautiful scenery if the user is relaxed."

[1163] This system can adapt to user emotions in real time and improve viewer engagement. In particular, in the case of advertising videos, dynamically changing content to match user emotions can maximize advertising effectiveness.

[1164] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1165] Step 1:

[1166] The server inputs music data and lyrics data. The server then quantifies the emotions in the lyrics and generates emotional data. Specifically, it uses natural language processing to analyze the main emotional elements of the lyrics and quantifies them, such as "70% joy, 30% sadness."

[1167] Step 2:

[1168] The server uses the quantified emotional data as input and automatically generates videos using a generative AI model. Based on the input emotional data, it constructs videos from highly relevant scenes. For example, if the emotional data is "70% joy, 30% sadness," it selects appropriate scenes (such as cherry blossoms in full bloom or a graduation ceremony) and generates a video.

[1169] Step 3:

[1170] The server reviews the generated video and makes any necessary adjustments, such as manually correcting scene timing or inappropriate parts of the generated video. As a result of the review, the final video file is generated.

[1171] Step 4:

[1172] The server uploads the final video to a database and distributes it to the user's device. Specifically, the video data is stored in the database along with meta information (song title, lyrics, and sentiment analysis results) and made accessible from the client device.

[1173] Step 5:

[1174] While the user is watching the advertising video, the device uses a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements. Specifically, the device's camera captures video data and the microphone captures audio data.

[1175] Step 6:

[1176] Based on the data acquired by the emotion engine, the user's emotions are analyzed and quantified emotional data is generated. Specifically, emotional data such as "80% excited, 20% relaxed" is generated from the user's facial expressions and tone of voice.

[1177] Step 7:

[1178] The generative AI model dynamically changes the video content based on emotional data acquired in real time. For example, if the user is excited, it will switch to scenes that emphasize the speed of cars, and if the user is relaxed, it will switch to scenes with beautiful scenery.

[1179] Step 8:

[1180] The device then plays the adjusted video in real time, ultimately displaying advertising videos that match the user's emotions, improving viewer engagement and maximizing advertising effectiveness.

[1181] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1182] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1183] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1184] [Fourth embodiment]

[1185] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1186] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1187] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1188] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1189] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1190] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1191] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1192] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1193] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1194] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1195] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1196] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1197] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1198] The present invention provides a system for automatically generating karaoke insert videos, and a specific embodiment thereof will be described below.

[1199] System Overview

[1200] The system is mainly composed of three components: a server, a device, and a user. The server automatically generates videos using a generative AI model and manages distribution to the device. The device provides the karaoke execution environment and an interface for users to enjoy karaoke. Users select a song and sing karaoke while visually enjoying the inserted video that matches the song.

[1201] Program processing overview

[1202] 1. The server collects video data from TV dramas and movies and corresponding music data, and prepares it as a training dataset for the generative AI model, which then learns the relationship between video scenes and music.

[1203] 2. The server inputs the new karaoke song and its lyrics, analyzes the emotions in the lyrics, and converts them into numerical values. For example, the main emotions in the lyrics are quantified as "joy 60%, sadness 40%."

[1204] 3. The server inputs the quantified emotional data into a generative AI model and automatically generates a video based on the emotional parameters. For example, for a song with a spring cherry blossom theme, a video including scenes of cherry blossoms in full bloom and students graduating will be generated.

[1205] 4. The server reviews the generated video and manually adjusts it as needed, for example by correcting the timing of certain scenes to optimize the synchronization between the video and the music.

[1206] 5. The server uploads the final video to a database and delivers it to the karaoke terminals, with the video data being transferred over a high-speed, stable connection.

[1207] 6. When the user selects a song, the device plays a video corresponding to that song. For example, if the user selects the song "Sakura," a pre-generated video of cherry blossoms in full bloom will be played.

[1208] 7. Users can visually enjoy the inserted video while singing the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[1209] Specific examples

[1210] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[1211] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[1212] 2. Server: The quantified emotional data is input into the generative AI model, and a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony that match the content of the lyrics.

[1213] 3. Server: Preview the generated video and manually fine-tune the timing of scenes, further refining the synchronization between the video and music.

[1214] 4. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[1215] 5. Device: When the user selects "Sakura," the video corresponding to the selected song will be played.

[1216] 6. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of graduation ceremonies.

[1217] In this way, this system can provide insert videos that match the worldview of the song when users enjoy karaoke, improving the visual entertainment value.In addition, by using a generative AI model, it is possible to significantly reduce the cost of video production and the burden of copyright fees.

[1218] The processing flow will be explained below.

[1219] Step 1:

[1220] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[1221] Step 2:

[1222] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and irrelevant data.

[1223] Step 3:

[1224] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and letting it learn the associations between videos and music based on the dataset.

[1225] Step 4:

[1226] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[1227] Step 5:

[1228] The server analyzes the emotions in the lyrics and converts them into numerical values, such as joy, anger, sadness, and pleasure. In this case, the result might be "70% joy, 30% sadness."

[1229] Step 6:

[1230] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[1231] Step 7:

[1232] The server reviews the generated video and manually adjusts it as needed, correcting scene timing or other inappropriate parts.

[1233] Step 8:

[1234] The server uploads the final video to a database, including meta information about the video data (such as song title, lyrics, and sentiment analysis results).

[1235] Step 9:

[1236] The terminal synchronizes with the karaoke system's database to receive new video data, ensuring that data transfer is fast and stable.

[1237] Step 10:

[1238] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[1239] Step 11:

[1240] The device plays the video corresponding to the song selected by the user. While the song "Sakura" is playing, scenes of cherry blossoms in full bloom and a graduation ceremony are displayed.

[1241] Step 12:

[1242] Users can visually enjoy the inserted video while singing along to the song, which helps them immerse themselves in the world of the song and improves the karaoke experience.

[1243] Example 1

[1244] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1245] In conventional karaoke systems, manually creating videos that match a song is time-consuming and costly, making it difficult to provide users with a satisfying entertainment experience. Furthermore, it is not possible to generate videos that match the emotions of the lyrics in real time, making it difficult to effectively convey the worldview of the song. To solve these issues, there is a need for a system that can efficiently and automatically generate videos that match a song and provide users with a high-quality karaoke experience.

[1246] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1247] In this invention, the server includes means for inputting music data and text data and quantifying emotions in the text data, means for using the quantified emotional data as input data and automatically generating videos using a generative AI model, means for reviewing the generated videos and fine-tuning them as necessary, means for uploading the final videos to a database and distributing them to a user terminal, and means for playing videos corresponding to the selected song when the user selects a song on the karaoke terminal. This makes it possible to efficiently and automatically generate videos suited to songs and provide users with a high-quality karaoke experience.

[1248] "Music data" is digital data of audio information that makes up music.

[1249] "Character data" is digital data that includes text information such as lyrics.

[1250] A "means for quantifying emotions" is a method or device that analyzes the content of character data and expresses its emotional elements as numerical values.

[1251] A "generative AI model" is an algorithm or program that uses machine learning and deep learning to automatically generate images.

[1252] "Means for automatically generating video" refers to a method or device that uses quantified emotional data and utilizes a generative AI model to create video.

[1253] The "reviewing means" refers to a method or device for checking the generated video and correcting it if necessary.

[1254] A "database" is an information storage system for storing and managing video and music data.

[1255] A "user terminal" is a device or apparatus that a user uses to enjoy karaoke.

[1256] A "karaoke terminal" is a device or apparatus dedicated to playing karaoke songs and displaying images.

[1257] The "video playback means" refers to a method or device for displaying video in accordance with the song selected on the karaoke terminal.

[1258] The present invention is a system for automatically generating images suitable for karaoke songs. A specific embodiment of this system will be described below.

[1259] The system mainly consists of three components: a server, a device, and a user. The server has a high-performance processor (e.g., an NVIDIA GPU) and large storage capacity, and uses machine learning frameworks such as TensorFlow or PyTorch to train and run the generative AI model. The device is an Android or iOS device, or a dedicated karaoke machine with built-in video playback software. The user uses a smartphone or a device in a karaoke booth.

[1260] The server first collects video data from TV dramas and movies and the corresponding music data. This creates a dataset for learning the association between video scenes and music. Specifically, the server obtains video and music data from the Internet or a dedicated database, and then formats them.

[1261] Next, the server inputs the new song and its text data. The text data includes information such as lyrics and title. This is analyzed using natural language processing (NLP) tools to quantify the emotions. For example, the lyrics of the song "Sakura" are quantified as "60% joy, 40% sadness." The analysis results are input into the generative AI model as prompt sentences.

[1262] Here is an example prompt: "Please analyze the emotions from the lyrics data of the specified song. Based on the analysis results, please generate an image that matches the lyrics. The emotional data is '70% joy, 30% sadness'."

[1263] The server uses the emotion data as input and automatically generates video using a generative AI model, which then creates a video based on the emotion data, generating scenes such as cherry blossoms in full bloom or a graduation ceremony.

[1264] The generated video is previewed on the server, and manual adjustments are made as needed, such as correcting the timing of specific scenes to optimize the synchronization between the video and music. Once the final adjustments are complete, the video is uploaded to the database.

[1265] The server delivers the videos uploaded to the database to the user's device. This delivery is done via a high-speed, stable connection. When the user selects a song, the device plays the video corresponding to that song in real time. For example, if the user selects "Sakura," a video of cherry blossoms in full bloom will be played as the intro begins.

[1266] Users can visually enjoy the video while singing the song. Because the video scenes are linked to the lyrics, users can immerse themselves deeply in the world of the song while singing. For example, while singing the song "Sakura," users can visually enjoy scenes of cherry blossoms in full bloom in spring or a graduation ceremony. In this way, users can be provided with a high-quality karaoke experience.

[1267] This system significantly reduces the cost of video production while providing users with video that matches the worldview of the song.

[1268] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1269] Step 1:

[1270] The server collects video data of TV dramas and movies and corresponding music data from the Internet or a dedicated database. This data is used as a training dataset for the generative AI model. The input data is the video and music data of TV dramas and movies, and the output data is a formatted training dataset. Specifically, the server obtains a scene from a movie and its background music and formats it into a dataset.

[1271] Step 2:

[1272] The server receives the newly added song and its lyrics data. It then uses a natural language processing (NLP) tool to quantify the emotion of the lyrics. The input data is the song and lyrics data, and the output data is the quantified emotion data. Specifically, the server analyzes the lyrics of the song "Sakura" and generates emotion data of "60% joy, 40% sadness."

[1273] Step 3:

[1274] The server converts the quantified emotional data into a prompt sentence and inputs it into the generative AI model. At this time, the generative AI model automatically generates a video based on the emotional parameters. The input data is the quantified emotional data, and the output data is the generated video. Specifically, the generated video includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[1275] Step 4:

[1276] The server reviews the generated footage and manually adjusts it as necessary. The input data is the generated footage, and the output data is the adjusted footage. Specifically, the server corrects the timing of specific scenes and optimizes the synchronization between the footage and the music.

[1277] Step 5:

[1278] The server uploads the final video to a database and distributes it to the user's device. The input data is the adjusted video, and the output data is the video stored in the database and the distributed video data. Specifically, the video data is stored in cloud storage and a link to it is sent to the device.

[1279] Step 6:

[1280] When the user selects a song, the device plays the video delivered from the server in real time. The input data is the user's song selection information and the delivered video data, and the output data is the video to be played. Specifically, when the user selects "Sakura," the intro starts and a video of cherry blossoms in full bloom is played at the same time.

[1281] Step 7:

[1282] Users visually enjoy the video while singing the song. The input data are karaoke songs and video data, and the output data is the user's karaoke experience. Specifically, users can sing "Sakura" while watching a touching graduation scene, enriching their karaoke experience.

[1283] (Application example 1)

[1284] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1285] There is a demand for combining music and video content to provide new content experiences that users can enjoy even more. However, creating visuals that match music is costly and requires a lot of effort. In addition, manual video production is time-consuming, and it is difficult to optimize the synchronization between music and video during the process. This issue needs to be resolved.

[1286] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1287] In this invention, the server includes means for inputting music data and lyric data and quantifying the emotion of the lyrics, means for automatically generating videos using a generative AI model using the quantified emotion data as input data, means for reviewing the generated videos and fine-tuning them as necessary, and means for playing videos corresponding to the selected songs when a user selects a song from a content distribution service. This improves the user experience, significantly reduces the cost and effort of video production, and enables highly accurate synchronization of music and video.

[1288] "Music data" refers to music information stored in digital format.

[1289] "Lyrics data" refers to the text information of the lyrics of a song stored in digital format.

[1290] "Emotional data" refers to the quantified information of emotional elements analyzed from lyrics.

[1291] A "generative AI model" refers to a machine learning model that uses artificial intelligence to generate new outputs from input data.

[1292] "Means for automatically generating videos" refers to the process of using an AI model to automatically create videos based on emotional data.

[1293] "Means of reviewing and fine-tuning as necessary" refers to the process of checking the generated video and manually correcting it as necessary.

[1294] "Content distribution service" refers to a service that provides digital content to users via the Internet.

[1295] "User terminal" refers to a device that a user uses to connect to the Internet, such as a smartphone or tablet.

[1296] This invention will explain a method for constructing a system in which music data and lyrics data are input, and related inserted animation is automatically generated and provided to the user.

[1297] System configuration

[1298] The system consists of the following hardware and software:

[1299] Server: Cloud server. Manages music data, analyzes lyrics sentiment, runs generative AI models, and manages generated videos.

[1300] Generative AI models: Use generative models such as GPT-4 and DALL-E.

[1301] Lyric analysis software: A text analysis tool that uses natural language processing technology.

[1302] Streaming software: Use HLS (HTTP Live Streaming) or similar.

[1303] User device: The device through which the user receives content, such as a smartphone or tablet.

[1304] Operational Overview

[1305] 1. The user selects a song

[1306] The user selects a song using the application of the content distribution service.

[1307] 2. The server retrieves the music data and lyrics data

[1308] Based on the song ID selected by the user, the server retrieves the song data and lyrics data.

[1309] 3. Lyrics Analysis

[1310] The server inputs the acquired lyrics data into lyrics analysis software and quantifies the emotions of the lyrics. For example, the main emotional elements of the lyrics are quantified as "joy 60% and sadness 40%."

[1311] 4. Visual Generation

[1312] The server inputs the quantified emotional data into a generative AI model to generate video visuals based on the emotions. For example, based on emotional data of "60% joy, 40% sadness," a video containing scenes of cherry blossoms in full bloom and students graduating might be generated.

[1313] 5. Review and fine-tune

[1314] The video generated on the server is previewed and manually tweaked as needed, correcting the timing of specific scenes and optimizing the synchronization of the video and music.

[1315] 6. Video distribution

[1316] The final fine-tuned video is delivered to the user's device using streaming software (such as HLS).

[1317] Specific examples

[1318] The prompt text and system behavior when the user selects the song "Spring Breeze" in the app are shown below.

[1319] Prompt statement

[1320] Visualize a spring landscape with emotional parameters of 70% hope and 30% joy.

[1321] System operation example

[1322] 1. When the user selects the song "Spring Breeze," the server retrieves the corresponding song data and lyrics data.

[1323] 2. Analyze the acquired lyrics data and quantify the emotional data (hope 70%, joy 30%).

[1324] 3. This emotional data is input into a generative AI model to generate images that express hope and joy, such as the fresh greenery of spring and bright skies.

[1325] 4. The generated video is reviewed on the server side, and the timing is fine-tuned to achieve optimal synchronization.

[1326] 5. The final video will be delivered to the user’s smartphone, where they can enjoy listening to the song “Spring Breeze.”

[1327] In this way, the synchronization between the music and the video can be improved, providing the user with a more attractive content experience.

[1328] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1329] Step 1:

[1330] The server receives an operation by the user to select a song and acquires the song ID.

[1331] Input: User-selected song ID

[1332] Output: Song ID is passed to the system

[1333] How it works: A user selects the song "Spring Breeze" on the app. The server detects the selection and retrieves the song ID (e.g., "12345").

[1334] Step 2:

[1335] The server uses the song ID to obtain the song data and lyrics data.

[1336] Input: Song ID

[1337] Output: Music data and lyrics data

[1338] Operation: The server accesses the database and retrieves the music data and lyrics data corresponding to the music ID "12345." At this time, the music data is retrieved in music file format, and the lyrics data is retrieved in text format.

[1339] Step 3:

[1340] The server analyzes the acquired lyrics data and quantifies the emotion of the lyrics.

[1341] Input: Lyrics data

[1342] Output: Emotion data (e.g., hope 70%, joy 30%)

[1343] How it works: The server inputs the lyrics data into natural language processing software, which analyzes the lyrics text. This analysis quantifies the emotional elements contained in the lyrics and generates emotional data, such as "70% hope, 30% joy."

[1344] Step 4:

[1345] The server inputs the quantified emotional data into a generative AI model and automatically generates videos based on the emotions.

[1346] Input: Emotion data

[1347] Output: Generated video

[1348] How it works: The server inputs the prompt "Please visualize a spring landscape with emotional parameters of 70% hope and 30% joy" into a generative AI model (e.g., DALL-E or GPT-4), and generates a video based on this. For example, a visual containing fresh spring greenery and a bright sky is generated.

[1349] Step 5:

[1350] The server reviews the generated video and makes any necessary adjustments.

[1351] Input: Generated video

[1352] Output: Final fine-tuned video

[1353] How it works: Preview the video generated on the server, check the timing and sequence of scenes, and if necessary, manually tweak specific scenes and timing to optimize the synchronization of music and visuals.

[1354] Step 6:

[1355] The server then delivers the final, finely tuned video to the user's device via streaming software.

[1356] Input: Final tweaked video

[1357] Output: Streaming video to user devices

[1358] How it works: The server uses HLS streaming software to encode the final video in real time and delivers it to the user's smartphone, tablet, or other device. Users can then watch the video corresponding to the song they selected on the app.

[1359] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1360] The present invention provides a system for automatically generating karaoke insert videos, which incorporates a technology that combines an emotion engine that recognizes the user's emotions. Specific embodiments of the technology are described below.

[1361] System configuration

[1362] The system mainly consists of the following components:

[1363] 1. Server

[1364] 2. User Device

[1365] 3. Emotion Engine

[1366] Program processing overview

[1367] 1. Server: Collects video and music data from TV dramas and movies and prepares it as a training dataset for the generative AI model, allowing the model to learn the relationship between video scenes and music.

[1368] 2. Server: A new karaoke song and its lyrics are input into the system, and the emotions in the lyrics are analyzed and quantified. For example, the main emotions in the lyrics are quantified as "joy 60% and sadness 40%."

[1369] 3. Server: The emotion engine recognizes the user's emotions in real time and feeds that emotion data back to the generative AI model. The emotion engine recognizes emotions from the user's facial expressions, tone of voice, body movements, etc.

[1370] 4. Server: Based on the emotional data obtained from the emotion engine and the quantified emotional data of the lyrics, the generative AI model automatically generates the optimal video. For example, if the user's emotions are "70% joy, 30% sadness," it will generate a video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[1371] 5. Server: Review the generated video and manually tweak it as needed, such as correcting scene timing or inappropriate parts.

[1372] 6. Server: Uploads the final video to the database and distributes it to the karaoke terminals, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[1373] 7. Terminal: Synchronizes with the karaoke system database and receives new video data, ensuring fast and stable data transfer.

[1374] 8. User: When a song is selected from the karaoke song list, a video corresponding to the song will be played, and the video will be adjusted in real time based on the user's emotions.

[1375] 9. Device: The device plays a video corresponding to the song selected by the user, and the emotion engine analyzes emotions in real time and adjusts the content of the video as needed.

[1376] 10. Users: They can visually enjoy the inserted video while singing the song. By adjusting the video based on the user's emotions, they can more easily immerse themselves in the world of the song, improving the karaoke experience.

[1377] Specific examples

[1378] For example, consider a case where a song titled "Sakura" is added to a karaoke service. The lyrics of this song include descriptions of the spring season and graduation scenes.

[1379] 1. Server: The song "Sakura" and its lyrics are input into the system, and the emotions in the lyrics are quantified. For example, they are quantified as "70% joy, 30% sadness."

[1380] 2. Emotion Engine: While the user is singing "Sakura," the system recognizes the user's emotions from their facial expressions, tone of voice, and body movements. For example, if the user is singing happily, it will be recognized as "80% joy, 20% sadness."

[1381] 3. Server: The emotion data recognized by the emotion engine is fed back to the AI ​​model, and combined with the emotional data from the lyrics, a video is generated that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[1382] 4. Server: Preview the generated video and manually fine-tune the timing of scenes to optimize the synchronization between the video and music.

[1383] 5. Server: The final video is uploaded to the database, and the video data corresponding to "Sakura" is distributed to the karaoke terminal.

[1384] 6. Device: When the user selects "Sakura," a video adjusted in real time based on the user's emotions is played.

[1385] 7. Users: While singing the song "Sakura," they can visually enjoy scenes of cherry blossoms in full bloom in spring and inserted videos of a graduation ceremony. In addition, the content of the video is adjusted in real time to match the user's emotions, allowing for a more emotionally immersive experience.

[1386] In this way, the system recognizes the user's emotions and uses that emotional data to adjust the video in real time, providing a visually enjoyable karaoke experience. Furthermore, the use of generative AI models can significantly reduce video production costs and copyright fees.

[1387] The processing flow will be explained below.

[1388] Step 1:

[1389] The server collects video data of TV dramas and movies and corresponding music data, including video clips for each scene and background music and scene music.

[1390] Step 2:

[1391] The server prepares the collected data as a training dataset for the generative AI model, and performs data cleaning to remove noise and inappropriate data.

[1392] Step 3:

[1393] The server inputs the prepared data into the generative AI model and starts the learning process, setting the model's initial parameters and training the model to recognize associations between videos and music based on the dataset.

[1394] Step 4:

[1395] The server inputs new karaoke songs and their lyrics data into the system, for example, the song "Sakura" and its lyrics into the system.

[1396] Step 5:

[1397] The server analyzes the emotions in the lyrics and quantifies emotions such as joy, anger, sadness, and happiness. In this case, the result might be "70% joy, 30% sadness."

[1398] Step 6:

[1399] The server inputs the quantified emotional data into a generative AI model, which then automatically generates videos based on the emotional parameters. The model then generates videos including scenes of cherry blossoms in full bloom and a graduation ceremony.

[1400] Step 7:

[1401] The server reviews the generated video and manually adjusts it to correct timing or inappropriate parts of the scene.

[1402] Step 8:

[1403] The server uploads the final video to a database, including meta information about the video data (song title, lyrics, and sentiment analysis results).

[1404] Step 9:

[1405] The terminal synchronizes with the karaoke system's database and receives new video data, ensuring fast and stable data transfer.

[1406] Step 10:

[1407] The user selects a song from the karaoke song list. When the user selects "Sakura," the video corresponding to that song is played.

[1408] Step 11:

[1409] The device plays a video corresponding to the song selected by the user. As the "Sakura" video plays, the emotion engine analyzes the user's emotions in real time.

[1410] Step 12:

[1411] The emotion engine recognizes the user's emotions in real time from their facial expressions, tone of voice, and body movements.

[1412] Step 13:

[1413] The server then feeds the emotion data recognized by the emotion engine back to the generative AI model, adjusting the video content in real time. For example, if the user is having fun, it adds a cheerful scene.

[1414] Step 14:

[1415] The device displays real-time updated video to the user, with the video content tailored to the user's current emotions.

[1416] Step 15:

[1417] Users can enjoy visually tailored videos while singing along, which helps users immerse themselves in the world of the song and enhances the karaoke experience.

[1418] In this way, by incorporating an emotion engine, it becomes possible to adjust the video content according to the user's real-time emotions, providing a more immersive karaoke experience.

[1419] Example 2

[1420] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1421] In conventional karaoke systems, the inserted video is fixed, making it impossible to dynamically adjust the video according to the user's emotions or singing style, resulting in a limited karaoke experience. In addition, creating fixed videos requires a great deal of effort and cost, making efficient operation difficult.

[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1423] In this invention, the server includes: a means for inputting music data and lyric data and quantifying the emotion of the lyrics; a means for automatically generating videos using a generative AI model using the quantified emotion data as input data; a means for reviewing the generated videos and fine-tuning them as necessary; an emotion recognition means for analyzing the user's facial expressions, tone of voice, and body movements in real time; a means for feeding back the analyzed emotion data in real time to the generative AI model and adjusting the content of the videos; a means for uploading the final videos to a database and distributing them to a user terminal; and a means for playing a video corresponding to a selected song when the user selects a song on the karaoke terminal. This enables dynamic video adjustment according to the user's emotion, providing a richer karaoke experience. It also improves video production efficiency and reduces operational costs.

[1424] "Music data" refers to audio data of music played on a karaoke system.

[1425] "Lyric data" is character information corresponding to a song, and is text data indicating the content of the lyrics.

[1426] An "emotion quantification means" is a method or system that performs a process of analyzing lyric content and assigning a numerical value to a particular emotional category.

[1427] A "generative AI model" is an artificial intelligence model that automatically generates videos from input data such as emotional data.

[1428] "Means for automatically generating videos" refers to the process of using a generative AI model to generate videos based on input data.

[1429] "Means for reviewing videos and fine-tuning them as needed" refers to the processes and tools used to review the content of the videos produced and adjust any inappropriate parts or timing.

[1430] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions, tone of voice, body movements, etc.

[1431] "Means for analyzing in real time" refers to a method for acquiring user emotion data in real time and executing an analysis process.

[1432] The "means for adjusting the content of the video" is a process for appropriately changing the scenes and content of the video being played based on the user's real-time emotional data.

[1433] The "means for uploading to a database" refers to the process or tool for storing the generated final video in a database such as cloud storage or a local server.

[1434] "Means for distributing to user terminals" refers to the process of transferring video data stored in the database to the karaoke terminals used by users.

[1435] The "means for playing video" refers to a function or process for playing video corresponding to a song selected on a user terminal using a display device or audio device.

[1436] The present invention is a system for automatically generating videos to be inserted into a karaoke system based on user emotion recognition. Specific embodiments of the system will be described below.

[1437] System Configuration

[1438] The system mainly consists of the following components:

[1439] 1. Server

[1440] 2. User Device

[1441] 3. Emotion Engine

[1442] Operation overview

[1443] Video and music data collection

[1444] Server: The server collects video and music data from TV dramas and movies from public databases on the Internet and from licensed content providers. This data is prepared as a dataset for training the generative AI model. Specifically, a Python script is used to download the data via an API and save it in storage. This process creates a dataset capable of learning a variety of emotional expressions and scenes.

[1445] Emotional analysis of new karaoke songs

[1446] Server: New karaoke songs and lyrics are input into the system, and natural language processing (NLP) techniques are used to quantify the emotions in the lyrics. This analysis uses emotion recognition libraries (e.g., NLTK, SpaCy). For example, the lyrics of the song "Sakura" are quantified as "70% joy, 30% sadness." This numerical data is used as input data for the generative AI model.

[1447] User Emotion Recognition

[1448] Emotion engine: Analyzes the user's facial expressions, vocal tone, and body movements in real time. This is done using a facial recognition camera, microphone, and motion sensor. The recognized emotional data is sent to a server. Specifically, facial expression analysis is performed using OpenCV, vocal tone analysis using a voice analysis tool, and body movements are analyzed using a motion sensor.

[1449] Optimal video generation

[1450] Server: Real-time emotional data obtained from the emotion engine and quantified emotional data from lyrics are input into the generative AI model. The generative AI model automatically generates optimal video scenes based on this data. For example, for "Sakura," it generates scenes of cherry blossoms in full bloom and a graduation ceremony. The prompt text in this case is as follows:

[1451] Generate an optimal insert video based on the lyrics of "Sakura." The emotional analysis of the lyrics reveals "70% joy, 30% sadness." The emotions recognized from the user's facial expressions, tone of voice, and body movements are "80% joy, 20% sadness." Based on this, create an inspiring video that includes scenes of cherry blossoms in full bloom and a graduation ceremony.

[1452] Video review and fine-tuning

[1453] Server: The server reviews the generated video and manually corrects any inaccuracies or timing adjustments. For example, they use video editing software such as Adobe Premiere Pro to optimize the synchronization between the video and the music.

[1454] Uploading video data

[1455] Server: Uploads the final video to a database and prepares it for distribution to user devices. This data includes meta information (song title, lyrics, sentiment analysis results). The server accesses the database and stores the video data in cloud storage or on a local server.

[1456] Data synchronization to karaoke terminal

[1457] User terminal: Periodically synchronizes with the karaoke system database to receive new video data. The terminal accesses the server and downloads data based on a set schedule.

[1458] Video playback and real-time adjustments

[1459] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. The emotion engine analyzes the user's emotions in real time and adjusts the video content as needed, allowing the user to enjoy a more emotionally immersive karaoke experience.

[1460] Through these steps, dynamic video adjustments based on user emotions are possible, providing a richer karaoke experience, while also improving video production efficiency and reducing operational costs.

[1461] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1462] Step 1:

[1463] Server: Collects video data from TV dramas and movies and music data, and prepares it as a training dataset for the generative AI model.

[1464] Input: Video and music data collected from public databases and content providers on the Internet.

[1465] Data processing: Using a Python script, the database is accessed via API, video and music data is downloaded, and stored in storage.

[1466] Output: The dataset used to train the generative AI model.

[1467] Step 2:

[1468] Server: New karaoke songs and their lyrics are input into the system, and the emotions of the lyrics are analyzed and quantified.

[1469] Input: Karaoke song and its lyrics data.

[1470] Data Arithmetic: Using an emotion recognition library (e.g., NLTK, SpaCy), parse the lyrics and assign numerical values ​​to major emotion categories, e.g., "70% joy, 30% sadness."

[1471] Output: Quantified lyrics emotion data.

[1472] Step 3:

[1473] Emotion engine: Identifies the user's facial expressions, tone of voice, and body movements in real time and sends the emotional data to the server.

[1474] Input: The user's facial expressions, tone of voice, and body movements.

[1475] Data calculation: Facial expressions, voice, and movements are captured using a facial recognition camera, microphone, and motion sensor, and then analyzed using OpenCV and voice analysis tools.

[1476] Output: Real-time sentiment data.

[1477] Step 4:

[1478] Server: Based on the emotional data from the emotion engine and the quantified emotional data from the lyrics, the generative AI model automatically generates optimal video scenes.

[1479] Input: Real-time emotion data from the emotion engine and quantified lyric emotion data.

[1480] Data computation: A generative AI model (e.g., GPT-4, DALL-E) generates video scenes based on prompts.

[1481] Output: The generated video scenes.

[1482] Step 5:

[1483] Server: Review the generated video and manually fine-tune any necessary parts.

[1484] Input: Generated video scenes.

[1485] Data processing: Use video editing software (e.g. Adobe Premiere Pro) to adjust inappropriate parts and timing.

[1486] Output: Final optimized video.

[1487] Step 6:

[1488] Server: Uploads the final video to the database and distributes it to the karaoke terminals.

[1489] Input: The final optimized video and its meta information (song title, lyrics, sentiment analysis results).

[1490] Data processing: Video data is stored in cloud storage or on a local server and prepared for distribution.

[1491] Output: Video data uploaded to the database.

[1492] Step 7:

[1493] Terminal: Synchronizes with the karaoke system database and receives new video data.

[1494] Input: New video data from the database.

[1495] Data processing: Through a periodic synchronization process, the server is accessed, video data is downloaded, and stored in the internal storage.

[1496] Output: New video data saved on your device.

[1497] Step 8:

[1498] Device: When a user selects a song from the karaoke song list, a video corresponding to that song is played. An emotion recognition engine adjusts the video content in real time as needed.

[1499] Input: User-selected song, real-time emotional data.

[1500] Data calculation: During video playback, the emotion recognition engine analyzes the user's emotions and adjusts the video content and timing accordingly.

[1501] Output: Real-time adjusted video playback provides a more emotionally immersive karaoke experience.

[1502] (Application example 2)

[1503] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1504] Modern advertising videos are unable to adapt their content in real time to the viewer's emotions, making it difficult to provide an optimal visual and emotional experience for each individual user, which limits the effectiveness of advertising and reduces viewer engagement.

[1505] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1506] In this invention, the server includes: means for inputting music data and lyric data and quantifying the emotion of the lyrics; means for automatically generating videos using a generative AI model using the quantified emotion data as input data; means for reviewing the generated videos and fine-tuning them as necessary; means for uploading the final videos to a database and distributing them to user terminals; means for acquiring the user's facial expressions, tone of voice, and movements in real time and analyzing their emotions; and means for the generative AI model to dynamically change the video content in real time based on the emotion data. This makes it possible to generate advertising videos that adapt to the user's emotions in real time, improve viewer engagement, and maximize advertising effectiveness.

[1507] "Music data" refers to the audio information of a musical work and its associated metadata.

[1508] "Lyric data" refers to text information about lyrics corresponding to a song.

[1509] The "emotion engine" is a function that analyzes emotions in real time from the user's facial expressions, voice, movements, etc.

[1510] "Numericalized emotion data" refers to data that expresses emotions as numerical values.

[1511] A "generative AI model" is an artificial intelligence model that generates new data or content based on input data.

[1512] "Automatic animation generation means" refers to a method or device for automatically generating animation using input data.

[1513] A "database" is a system for efficiently storing, managing, and retrieving data.

[1514] A "user terminal" is an electronic device that is directly operated by a user.

[1515] "Tweaking tool" refers to a method or device for manually adjusting generated content.

[1516] "Facial expression recognition" is a technology that analyzes a user's facial expressions and identifies their emotions based on them.

[1517] "Tone of voice" refers to characteristics such as pitch, intensity, and speed of voice, and is an element that conveys the speaker's emotions and intentions.

[1518] "Real-time acquisition" refers to the acquisition and processing of data immediately, without delay.

[1519] "Dynamic change" means that a system or content changes automatically over time or in response to events.

[1520] An "advertising video" is video content created to promote a particular product or service.

[1521] "Viewer's emotions" refer to the emotions and feelings felt by a user watching video content.

[1522] The present invention relates to a system for dynamically generating and adjusting advertising videos in real time based on user emotions. Specific embodiments of the system are described below.

[1523] System configuration

[1524] The system mainly consists of the following components:

[1525] 1. Server

[1526] 2. User Device

[1527] 3. Emotion Engine

[1528] Program processing overview

[1529] Hardware and Software Configuration

[1530] Hardware:

[1531] Smartphone camera

[1532] Smartphone microphone

[1533] software:

[1534] EmotionEngine (emotion analysis engine)

[1535] Video Generator (generative AI model)

[1536] Database management systems (e.g., SQLite)

[1537] System Operation

[1538] 1. The server first inputs the music data and lyrics data and quantifies the emotion of the lyrics.

[1539] 2. Using the quantified emotional data as input data, a generative AI model is used to automatically generate videos.

[1540] 3. Review the generated video on the server side and make any necessary adjustments.

[1541] 4. Upload the final video to the database and distribute it to the user's device.

[1542] 5. While the user is watching the advertising video, the user's device will use a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements.

[1543] 6. The emotion engine analyzes the user's emotions based on the acquired data and generates quantified emotion data.

[1544] 7. The generative AI model dynamically changes video content based on emotional data captured in real time.

[1545] 8. The user device plays the adjusted video in real time.

[1546] Specific examples

[1547] For example, consider an advertising video for a "new sports car." This advertisement dynamically changes to emphasize the speed of the car when the user is excited, and to show scenes of the car traveling through beautiful scenery when the user is relaxed. Specific examples of prompt sentences are as follows:

[1548] Example prompt:

[1549] "For an ad for a new sports car, generate a video that shows a car speeding along if the user is excited, or a video that shows a car driving through beautiful scenery if the user is relaxed."

[1550] This system can adapt to user emotions in real time and improve viewer engagement. In particular, in the case of advertising videos, dynamically changing content to match user emotions can maximize advertising effectiveness.

[1551] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1552] Step 1:

[1553] The server inputs music data and lyrics data. The server then quantifies the emotions in the lyrics and generates emotional data. Specifically, it uses natural language processing to analyze the main emotional elements of the lyrics and quantifies them, such as "70% joy, 30% sadness."

[1554] Step 2:

[1555] The server uses the quantified emotional data as input and automatically generates videos using a generative AI model. Based on the input emotional data, it constructs videos from highly relevant scenes. For example, if the emotional data is "70% joy, 30% sadness," it selects appropriate scenes (such as cherry blossoms in full bloom or a graduation ceremony) and generates a video.

[1556] Step 3:

[1557] The server reviews the generated video and makes any necessary adjustments, such as manually correcting scene timing or inappropriate parts of the generated video. As a result of the review, the final video file is generated.

[1558] Step 4:

[1559] The server uploads the final video to a database and distributes it to the user's device. Specifically, the video data is stored in the database along with meta information (song title, lyrics, and sentiment analysis results) and made accessible from the client device.

[1560] Step 5:

[1561] While the user is watching the advertising video, the device uses a camera and microphone to capture real-time data such as the user's facial expressions, tone of voice, and movements. Specifically, the device's camera captures video data and the microphone captures audio data.

[1562] Step 6:

[1563] Based on the data acquired by the emotion engine, the user's emotions are analyzed and quantified emotional data is generated. Specifically, emotional data such as "80% excited, 20% relaxed" is generated from the user's facial expressions and tone of voice.

[1564] Step 7:

[1565] The generative AI model dynamically changes the video content based on emotional data acquired in real time. For example, if the user is excited, it will switch to scenes that emphasize the speed of cars, and if the user is relaxed, it will switch to scenes with beautiful scenery.

[1566] Step 8:

[1567] The device then plays the adjusted video in real time, ultimately displaying advertising videos that match the user's emotions, improving viewer engagement and maximizing advertising effectiveness.

[1568] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1569] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1570] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1571] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1572] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1573] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1574] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1575] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1576] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1577] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1578] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1579] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1580] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1581] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1582] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1583] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1584] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1585] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1586] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1587] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1588] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1589] The following is further disclosed regarding the above embodiment.

[1590] (Claim 1)

[1591] A means to input music data and lyrics data and quantify the emotion of the lyrics,

[1592] A means to automatically generate videos using a generative AI model using quantified emotional data as input data;

[1593] A means to review the generated video and fine-tune it as needed;

[1594] A means for uploading the final video to a database and distributing it to user devices;

[1595] means for playing a video corresponding to a song selected by a user on a karaoke terminal;

[1596] A system including:

[1597] (Claim 2)

[1598] The system of claim 1, wherein the generative AI model is trained based on video data and music data from television dramas and movies.

[1599] (Claim 3)

[1600] The system of claim 1, which quantifies the emotions of lyrics into four categories: joy, anger, sadness, and happiness.

[1601] "Example 1"

[1602] (Claim 1)

[1603] A means for inputting music data and text data and quantifying the emotions of the text data;

[1604] A means to automatically generate video using a generative AI model using quantified emotional data as input data;

[1605] A means to review the resulting footage and fine-tune it as needed;

[1606] A means for uploading the final video to a database and distributing it to user terminals;

[1607] means for playing back a video corresponding to a song selected by a user on a karaoke terminal;

[1608] A system including:

[1609] (Claim 2)

[1610] The system of claim 1, wherein the generative AI model is trained based on video data and audio data from dramas and movies.

[1611] (Claim 3)

[1612] 10. The system of claim 1, wherein emotions in the text data are quantified into categories.

[1613] "Application Example 1"

[1614] (Claim 1)

[1615] A means to input music data and lyrics data and quantify the emotion of the lyrics,

[1616] A means to automatically generate videos using a generative AI model using quantified emotional data as input data;

[1617] A means to review the generated video and fine-tune it as needed;

[1618] A means for uploading the final video to a database and distributing it to user devices;

[1619] means for playing a video corresponding to a song selected by a user through a content distribution service;

[1620] A system including:

[1621] (Claim 2)

[1622] 10. The system of claim 1, wherein the generative AI model is trained based on video data and music data.

[1623] (Claim 3)

[1624] 10. The system of claim 1, wherein the system quantifies the emotion of lyrics into a plurality of emotion categories.

[1625] "Example 2: Combining Emotion Engines"

[1626] (Claim 1)

[1627] A means to input music data and lyrics data and quantify the emotion of the lyrics,

[1628] A means to automatically generate videos using a generative AI model using quantified emotional data as input data;

[1629] A means to review the generated video and fine-tune it as needed;

[1630] A means for uploading the final video to a database and distributing it to user devices;

[1631] An emotion recognition means that analyzes the user's facial expressions, tone of voice, and body movements in real time;

[1632] A method to adjust the content of videos by feeding back real-time analyzed emotional data to the generative AI model, and

[1633] means for playing a video corresponding to a song selected by a user on a karaoke terminal;

[1634] A system including:

[1635] (Claim 2)

[1636] 10. The system of claim 1, wherein the generative AI model is trained based on video data and audio data.

[1637] (Claim 3)

[1638] 10. The system of claim 1, wherein the system quantifies the emotion of lyrics into a plurality of categories.

[1639] "Application example 2 when combining emotion engines"

[1640] (Claim 1)

[1641] A means to input music data and lyrics data and quantify the emotion of the lyrics,

[1642] A means to automatically generate videos using a generative AI model using quantified emotional data as input data;

[1643] A means to review the generated video and fine-tune it as needed;

[1644] A means for uploading the final video to a database and distributing it to user devices;

[1645] means for playing a video corresponding to a song selected by a user on a karaoke terminal;

[1646] A means of capturing the user's facial expressions, tone of voice, and movements in real time and analyzing their emotions;

[1647] A means for a generative AI model to dynamically change video content in real time based on emotion data,

[1648] A system including:

[1649] (Claim 2)

[1650] The system of claim 1, wherein the generative AI model is trained based on video data and music data from television dramas and movies.

[1651] (Claim 3)

[1652] The system of claim 1, which quantifies the emotions of lyrics into four categories: joy, anger, sadness, and happiness. [Explanation of symbols]

[1653] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means to input music data and lyrics data and quantify the emotion of the lyrics, A means to automatically generate videos using a generative AI model using quantified emotional data as input data; A means to review the generated video and fine-tune it as needed; A means for uploading the final video to a database and distributing it to user devices; means for playing a video corresponding to a song selected by a user on a karaoke terminal; A system including:

2. The system of claim 1, wherein the generative AI model is trained based on video data and music data from television dramas and movies.

3. The system according to claim 1, wherein the emotions of lyrics are quantified into four categories: joy, anger, sadness, and pleasure.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A