system

A system using generative AI and NFTs to replicate voice actors' voices and distribute revenue addresses quality decline and income instability, ensuring authentic voice reproduction and continuous model improvement.

JP2026034004APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137125
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In traditional anime production, the absence of a voice actor due to illness or retirement leads to a decline in quality and instability in voice actors' income, with no effective method to maintain voice authenticity and collect viewer feedback for model improvement.

Method used

A system utilizing generative AI to replicate voice actors' voices, assigning NFTs for ownership, and a revenue distribution mechanism, combined with viewer feedback for continuous model improvement.

Benefits of technology

Maintains anime quality and stabilizes voice actors' income by ensuring authentic voice replication and revenue sharing, while enabling AI model refinement through viewer feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034004000001_ABST
    Figure 2026034004000001_ABST
Patent Text Reader

Abstract

To provide a system which duplicates voices of voice actors in a generation AI and stabilizes profits of the voice actors while maintaining qualities of works.SOLUTION: A generated AI model training unit configured to train a generated AI model based on analyzed feature information, a generated voice evaluation unit configured to evaluate qualities of voices generated by the generated AI model, a voice NFT attaching unit configured to generate NFTs for original voices of voice actors and register the NFTs in blockchains, a generated voice downloading unit configured to download the generated voices, a revenue sharing unit configured to record revenues and return the revenues to the voice actors, and an audience feedback collecting unit configured to collect and analyze audience feedback.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In traditional anime production, if a voice actor is unable to perform due to illness or retirement, the character's voice must be changed, which can cause discomfort to viewers and lead to a decline in the quality of the work. Furthermore, voice actors' income is unstable, leaving limited means for them to secure a continuous income. To solve these issues, a system is needed that uses generative AI to replicate the voice of a voice actor, maintaining the quality of the work while stabilizing the voice actor's income. [Means for solving the problem]

[0005] The present invention provides a system that solves the above-mentioned problems by the following means. First, it includes a voice analysis means that analyzes a voice actor's voice data and trains a generative AI model based on the analysis data. Next, it includes a generated voice evaluation means that evaluates the quality of the voice generated by the generative AI model and improves the AI ​​model based on the evaluation results. It also includes an NFT assignment means for assigning an NFT to the voice actor's original voice data and registering the NFT on the blockchain. It also includes a generated voice download means that allows animation production companies to easily use the generated voice, and a revenue distribution means that records revenue from the produced work and returns a portion to the voice actor. Finally, it includes a viewer feedback collection means that collects feedback from viewers and uses it to improve the AI ​​model. This allows the voice actor's voice quality to be maintained while ensuring stable revenue.

[0006] "Voice analysis means" refers to a device or program that has the function of analyzing the voice data recorded by the voice actor and extracting voice characteristics such as waveform, spectrum, pitch, volume, and intonation.

[0007] "Training means for generative AI model" refers to a device or program that trains a generative AI model based on analyzed voice feature data and carries out a learning process to reproduce the unique voice quality of a voice actor.

[0008] "Means for evaluating generated speech" refers to a device or program that has the function of evaluating the quality of speech generated by a generative AI model and measuring the similarity and naturalness by comparing it with the original.

[0009] A "means for assigning NFTs to audio data" is a device or program that has the function of generating NFTs for a voice actor's original audio data and registering the NFTs on a blockchain, thereby guaranteeing ownership and authenticity.

[0010] A "means for downloading generated audio" is a device or program that has the function of allowing an animation production company to download generated audio data and original audio NFTs from a server through an appropriate interface.

[0011] A "revenue distribution means" is a device or program that has the function of recording the revenue from the produced anime work and carrying out the calculation and payment process to return a portion of that revenue to the voice actors.

[0012] A "means for collecting feedback from viewers" is a device or program that has the function of collecting feedback from viewers who have watched the completed anime work and using that feedback to help improve the generative AI model. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] overview

[0035] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns NFTs to them to preserve the value of the originals. This aims to maintain the quality of anime works even when voice actors are unable to perform, and to stabilize voice actors' earnings.

[0036] Overall system configuration

[0037] The system of the present invention comprises the following elements:

[0038] 1. Voice analysis method: Analyze the voice actor's voice data.

[0039] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0040] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0041] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0042] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0043] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0044] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0045] Explanation of system processing

[0046] 1. Collection and analysis of audio data

[0047] Subject: Terminal

[0048] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[0049] The device uploads the recorded audio data to the server.

[0050] Subject: Server

[0051] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[0052] 2. Training a generative AI model

[0053] Subject: Server

[0054] The server uses the analyzed audio feature data to train a generative AI model.

[0055] This AI model is designed to reproduce the unique vocal qualities of voice actors.

[0056] 3. Evaluation of generated speech

[0057] Subject: Server

[0058] The server generates new voices using generative AI and tests them to evaluate their quality.

[0059] The model is repeatedly improved based on the evaluation results.

[0060] 4. NFT assignment to original audio

[0061] Subject: Server

[0062] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[0063] Registered NFTs prove ownership and authenticity of audio data.

[0064] 5. Audio Use and Revenue Sharing

[0065] Subject: Terminal (animation production company)

[0066] The device downloads the generated audio data and NFT from the server.

[0067] Use synthetic audio in animation production and release the finished work.

[0068] The revenue of the work is recorded and sent to the server.

[0069] Subject: Server

[0070] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[0071] 6. Quality check and feedback

[0072] Subject: User (viewer)

[0073] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[0074] Subject: Server

[0075] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[0076] Specific examples

[0077] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This system eliminates the decline in quality and instability of revenue that can occur when voice actors are replaced.

[0078] The processing flow will be explained below.

[0079] Step 1:

[0080] Audio data collection

[0081] Subject: Terminal

[0082] The device records voice actors' voices in high quality, including the characters' lines and various emotional expressions.

[0083] Step 2:

[0084] Uploading audio data

[0085] Subject: Terminal

[0086] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure secure transmission.

[0087] Step 3:

[0088] Analysis of audio data

[0089] Subject: Server

[0090] The server analyzes the received audio data. This analysis involves extracting features such as the waveform, spectrum, pitch, volume, and intonation of the audio. The analysis results are stored in a database.

[0091] Step 4:

[0092] Training generative AI models

[0093] Subject: Server

[0094] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[0095] Step 5:

[0096] Generating synthetic speech

[0097] Subject: Server

[0098] The server uses a trained generative AI model to generate new voice samples, which are then compared to the original voice.

[0099] Step 6:

[0100] Evaluation of generated speech

[0101] Subject: Server

[0102] The server evaluates the quality of the generated voice, including voice similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[0103] Step 7:

[0104] NFT granting

[0105] Subject: Server

[0106] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the voice data.

[0107] Step 8:

[0108] Download generated audio

[0109] Subject: Terminal (animation production company)

[0110] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[0111] Step 9:

[0112] Anime production and revenue records

[0113] Subject: Terminal

[0114] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[0115] Step 10:

[0116] Revenue sharing

[0117] Subject: Server

[0118] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payments.

[0119] Step 11:

[0120] Gathering feedback from viewers

[0121] Subject: User

[0122] Users can watch the completed animation and provide feedback to the platform on the quality of the generated audio and their impressions.

[0123] Step 12:

[0124] Feedback analysis and AI model improvement

[0125] Subject: Server

[0126] The server collects and analyzes feedback from viewers, and the analysis results are used to further refine the AI ​​model, thereby improving the quality of the generated voice in future iterations.

[0127] Example 1

[0128] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0129] If a voice actor is unable to provide voice over for illness or other reasons, the quality of the anime work declines and revenue becomes unstable. Furthermore, there is no way to guarantee the authenticity of the original voice over, which creates the risk of counterfeiting or unauthorized use. Therefore, a method is needed to effectively collect viewer feedback and improve generative AI models.

[0130] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0131] In this invention, the server includes a voice analysis means for analyzing voice data and extracting feature data; a generative AI model training means for training a generative AI model based on the analyzed feature data; a generated voice evaluation means for generating new voices using the generative AI model and evaluating their quality; a voice data NFT assignment means for assigning NFTs to the voice actor's original voice data and registering it on the blockchain; a revenue distribution means for returning revenue to the voice actor based on revenue data; and a feedback collection means for collecting feedback from viewers and using it to improve the generative AI model. This enables the production of high-quality anime works even when voice actors are unavailable, improving revenue stability. It also ensures the authenticity of the original voice and enables continuous improvement of the AI ​​model through feedback.

[0132] 1. "Audio data" means data that is a digital recording of the sounds and lines spoken by a voice actor.

[0133] 2. "Terminal means" means a hardware device or software tool for recording or uploading audio data.

[0134] 3. "Server" refers to a computer system that performs processes such as analyzing voice data, evaluating generated voice data, and assigning NFTs.

[0135] 4. "Audio analysis means" means technologies or algorithms that process uploaded audio data to extract characteristics such as waveform, spectrum, pitch, volume, and intonation.

[0136] 5. A "generative AI model" is a machine learning model that generates new speech based on speech data.

[0137] 6. "Generative AI model training means" means the process of optimizing and training a generative AI model using analyzed speech feature data.

[0138] 7. “Generative Speech Evaluation Measures” means tests or criteria for assessing the quality of speech generated by a generative AI model.

[0139] 8. "NFT" stands for "Non-Fungible Token" and is a token used to prove ownership and authenticity of digital assets.

[0140] 9. "Audio data NFT assignment method" is a technology that generates NFTs for audio data and registers them on the blockchain.

[0141] 10. "Revenue sharing mechanism" is a system for recording revenue from anime works and distributing that revenue fairly to voice actors.

[0142] 11. “Feedback collection methods” are technologies and methods used to collect audience ratings and feedback and use it to improve the generative AI model.

[0143] MODE FOR CARRYING OUT THE INVENTION

[0144] This invention is a system that uses generative AI technology to replicate the voice of a voice actor and assigns an NFT to preserve the value of the original. Specific embodiments for implementing this system are described below.

[0145] 1. Hardware Configuration

[0146] Terminal

[0147] Use high-quality dedicated microphones and studio equipment to record audio data, and a computer or dedicated device to capture the recorded audio data in digital format (e.g., WAV files) and upload it to a server via an internet connection.

[0148] server

[0149] It uses a high-performance server that analyzes voice data, trains generative AI models, evaluates generated voices, awards NFTs, distributes revenue, and collects feedback. Specifically, it refers to a server with the computing resources to run libraries such as Python, TENSORFLOW (registered trademark), and PyTorch.

[0150] 2. Software Configuration

[0151] Voice analysis methods

[0152] We use LibROSA or other audio analysis libraries to extract features from the audio data, such as waveform, spectrum, pitch, volume, and intonation, and store these in a database.

[0153] Generative AI model training tools

[0154] To train the generative AI model, we use Python and machine learning libraries such as TensorFlow and PyTorch. Using the analyzed audio feature data, we optimize the model (e.g., WaveNet, Tacotron2) to reproduce the unique voice quality of the voice actor.

[0155] Generated speech evaluation means

[0156] To evaluate the quality of the generated speech, we evaluate the performance of the model using speech evaluation criteria such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[0157] Audio data NFT granting method

[0158] To assign an NFT to audio data, a smart contract creation tool (e.g., Solidity) is used and registered on the Ethereum blockchain to prove the authenticity and ownership of the audio data.

[0159] Revenue sharing method

[0160] Revenue data is collected through online payment systems (e.g., PayPal, Stripe), and revenue is distributed to voice actors based on that data. Revenue distribution is calculated using automated scripts and programs on the server.

[0161] Feedback collection methods

[0162] A dedicated feedback form and application are used to collect feedback from viewers, and the collected data is used as an evaluation index for the generative AI model, helping to improve the model.

[0163] Specific examples

[0164] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. The original voice is assigned an NFT to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This process prevents a decline in quality due to the voice actor's absence.

[0165] Prompt Sentence Examples

[0166] "As a first step, please record voice actor A's past voice data and upload it to our server. Next, we will use the server to analyze this voice data and train a generative AI model. Once the model is trained, we will evaluate the quality of the generated voice and make improvements if necessary. Finally, please attach an NFT to the original voice and download the generated voice for use."

[0167] This invention makes it possible to produce high-quality anime even when voice actors are unavailable, improving revenue stability and enabling continuous improvement of AI models based on viewer feedback.

[0168] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0169] Step 1: Record and upload audio data

[0170] Subject: Terminal

[0171] The device uses dedicated microphones and studio equipment to record high-quality voice data from voice actors. The recorded voice data (e.g., in WAV format) includes the character's lines and emotional expressions.

[0172] Once recording is complete, the device uploads the audio data to a server over the internet using a secure file transfer protocol (e.g., SFTP).

[0173] Input: Voice actor's voice data

[0174] Output: Upload audio data to the server

[0175] Step 2: Analyzing the audio data

[0176] Subject: Server

[0177] The server receives the uploaded audio data and uses an audio analysis library such as LibROSA to extract features such as waveform, spectrum, pitch, volume, and intonation.

[0178] The analyzed voice feature data is stored in a database.

[0179] Input: Uploaded audio data

[0180] Output: Analyzed audio feature data

[0181] Step 3: Training the generative AI model

[0182] Subject: Server

[0183] The server trains a generative AI model based on the analyzed voice feature data, using Python and machine learning libraries such as TensorFlow and PyTorch. The generative AI model employs a multilayer neural network-based model such as WaveNet or Tacotron2.

[0184] The model is trained to replicate the voice actor's unique vocal timbre and speaking style.

[0185] Input: Analyzed audio feature data

[0186] Output: A trained generative AI model

[0187] Step 4: Evaluate and improve the generated speech

[0188] Subject: Server

[0189] The server generates test audio using a trained generative AI model, which is then evaluated using audio metrics such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[0190] Based on the evaluation results, improvements are made repeatedly by adjusting the model's hyperparameters and retraining with additional data.

[0191] Input: Trained generative AI model, test audio data

[0192] Output: Evaluation results, improved generative AI model

[0193] Step 5: Adding an NFT to the original audio

[0194] Subject: Server

[0195] The server generates an NFT for the original audio data, using a smart contract to prove ownership and authenticity of the audio data.

[0196] The generated NFT is registered on the Ethereum blockchain and acts as a digital certificate.

[0197] Input: Original audio data

[0198] Output: NFT, registered on the blockchain

[0199] Step 6: Download and use the audio data

[0200] Subject: Terminal (animation production company)

[0201] The device downloads the generated audio data and NFT from the server, which is done securely using SFTP.

[0202] The animation production company will use the downloaded generated audio to create and release an animated work.

[0203] Input: Generated audio data, NFT

[0204] Output: Finished animation

[0205] Step 7: Revenue sharing

[0206] Subject: Terminal (animation production company)

[0207] The terminal records the revenue generated by the published anime works and transmits the revenue data to the server, which collects the data through online payment systems (e.g., PayPal, Stripe).

[0208] Input: Revenue data for completed anime works

[0209] Output: Send revenue data to the server

[0210] Subject: Server

[0211] The server distributes revenue to voice actors based on the received revenue data. Revenue distribution calculations are performed using automated scripts or programs based on pre-set percentages.

[0212] Input: Revenue data for anime works

[0213] Output: Revenue return to voice actors

[0214] Step 8: Gather viewer feedback and refine the model

[0215] Subject: User (viewer)

[0216] After watching the anime, users can provide their impressions and audio quality ratings using a dedicated feedback form or app.

[0217] Input: Viewer feedback

[0218] Output: Feedback data

[0219] Subject: Server

[0220] The server collects feedback from viewers and analyzes it as an evaluation index for machine learning. Based on the results of this analysis, the generative AI model is further improved.

[0221] Input: Feedback data

[0222] Output: An improved generative AI model

[0223] (Application example 1)

[0224] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0225] Conventional voice generation systems for voice actors have had problems maintaining the quality of voice data and stable revenue when voice actors are unable to physically participate in recording. Furthermore, there is no way to guarantee the authenticity or ownership of the generated voice data, making it difficult to effectively collect feedback from viewers and use it to improve AI models. The present invention aims to solve these problems by stably generating voice data for voice actors, ensuring revenue, and effectively utilizing feedback from viewers.

[0226] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0227] In this invention, the server includes a voice analysis means, a training means for a generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, a means for listening to the voice data through a listening application installed on a smartphone, a means for purchasing NFTs via the listening application, and a means for collecting feedback information via the listening application. This allows the quality of the generated voice and stability of revenue to be maintained even in situations where the voice actor cannot participate in recording, and enables the AI ​​model to be improved based on feedback from viewers.

[0228] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[0229] A "training method for a generative AI model" is a method for training an AI model based on analyzed voice data.

[0230] The "means for evaluating generated speech" is a means for evaluating and testing the quality of generated speech.

[0231] "Method of assigning NFTs to audio data" is a method of assigning NFTs to voice actor audio data to prove its authenticity and ownership.

[0232] The "means for downloading generated voice" is a means for making the generated voice data downloadable.

[0233] The "profit distribution means" is a means for appropriately distributing the profits from the content in which the generated audio is used.

[0234] "Means for collecting feedback from viewers" refers to means for collecting opinions and impressions from viewers.

[0235] "Means for listening to audio data through a listening application installed on a smartphone" refers to means for listening to audio data through an application installed on a smartphone.

[0236] "Means for purchasing NFTs via a viewing application" refers to means for purchasing NFTs using a viewing application.

[0237] The "means for collecting feedback information via a viewing application" refers to a means for collecting feedback information from viewers through a viewing application.

[0238] System configuration

[0239] The system of the present invention comprises the following elements:

[0240] 1. Voice analysis method: Analyze the voice actor's voice data.

[0241] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0242] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0243] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0244] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0245] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0246] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0247] 8. Means for listening to audio data through a listening application installed on a smartphone: A smartphone application is used as a means for a user to listen to audio data.

[0248] 9. Means for purchasing NFTs via the viewing application: A means for users to purchase NFTs via the viewing application.

[0249] 10. Means of collecting feedback information via a viewing application: Means of collecting viewer feedback information through a viewing application.

[0250] Explanation of program processing

[0251] 1. Collection and analysis of audio data

[0252] Subject: Terminal

[0253] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[0254] The device uploads the recorded audio data to the server.

[0255] Subject: Server

[0256] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and modulation, and stores the data in a database.

[0257] 2. Training a generative AI model

[0258] Subject: Server

[0259] The server uses the analyzed audio feature data to train a generative AI model, which is built using libraries such as TensorFlow and PyTorch.

[0260] 3. Evaluation of generated speech

[0261] Subject: Server

[0262] The server generates new voices using the generative AI and performs tests to evaluate their quality. The model is then repeatedly improved based on the evaluation results. A Python evaluation algorithm is used for the evaluation.

[0263] 4. NFT assignment to original audio

[0264] Subject: Server

[0265] The server generates an NFT using Ethereum for the voice actor's original voice data and registers it on the blockchain via the OpenSea API.

[0266] 5. Audio Use and Revenue Sharing

[0267] Subject: Terminal (animation production company)

[0268] The device downloads the generated audio data and NFTs from the server, uses the generated audio in animation production, and publishes the finished work. The revenue from the work is recorded and sent to the server.

[0269] Subject: Server

[0270] The server receives the revenue data and returns the revenue to the voice actors based on a preset percentage. AWS (registered trademark) is used as the revenue management system.

[0271] 6. Quality check and feedback

[0272] Subject: User (viewer)

[0273] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[0274] Subject: Server

[0275] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[0276] Examples of specific examples and prompts

[0277] For example, we will show a specific example using the smartphone app "Spoken NFT."

[0278] 1. App launch: The user launches the "Spoken NFT" app and searches for the generated voice of their favorite voice actor.

[0279] 2. Listen to the audio: Select the generated audio from the search results and listen to it on your smartphone.

[0280] 3. Purchase NFT: Purchase an NFT for the generated audio you like and retain ownership on the blockchain.

[0281] 4. Provide feedback: After listening, send feedback about the generated audio through the app.

[0282] Here are some example prompts to input to a generative AI model:

[0283] "Generate lines for character A with the characteristics of voice actor XX:

[0284] "From today onwards, you are a part of our team!"

[0285] Emotion: Joy

[0286] Based on this prompt, the AI ​​will reproduce the voice quality of voice actor XX and the characteristics of character A, and generate voice that also takes emotional expression into account.

[0287] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0288] Step 1: Collecting audio data

[0289] Subject: Terminal

[0290] The device records high-quality voice data from the voice actor. This voice includes the character's lines and emotional expressions. The device then uploads this voice data to the server. The input is the recorded voice data, and the output is the audio file uploaded to the server.

[0291] Step 2: Analyzing the audio data

[0292] Subject: Server

[0293] The server analyzes the received audio data by breaking it down into features such as waveform, spectrum, pitch, volume, and modulation. The analysis results are stored in a database. The input is the uploaded audio file, and the output is the analyzed audio feature data.

[0294] Step 3: Training the generative AI model

[0295] Subject: Server

[0296] The server uses the analyzed voice feature data to train a generative AI model using TensorFlow or PyTorch. During the training process, data processing and calculations are performed to update the AI ​​model. The input is the analyzed voice feature data, and the output is a trained generative AI model.

[0297] Step 4: Generate and evaluate synthetic speech

[0298] Subject: Server

[0299] The server generates new speech data using a trained generative AI model, then evaluates the quality of the generated speech and refines the AI ​​model as needed. The input is the trained generative AI model and a text prompt, and the output is the generated new speech data.

[0300] Step 5: Adding an NFT to the audio data

[0301] Subject: Server

[0302] The server generates an NFT for the generated audio data using Ethereum and registers it on the blockchain via the OpenSea API. The input is the generated audio data, and the output is the audio data with the NFT attached and a record of its ownership.

[0303] Step 6: Download the generated audio

[0304] Subject: Terminal (animation production company)

[0305] The device downloads the generated audio data and NFT from the server. The input is the audio data with the NFT attached, and the output is the downloaded audio file.

[0306] Step 7: Revenue sharing

[0307] Subject: Server

[0308] The server receives revenue data for the work and distributes the revenue to the voice actors based on a preset percentage. The input is revenue data, and the output is distributed revenue information.

[0309] Step 8: Gather feedback from your audience

[0310] Subject: User (viewer)

[0311] After listening to the generated speech through a smartphone application, the user sends feedback. The input is the feedback content, and the output is the feedback data sent to the server.

[0312] Step 9: Improve the AI ​​model with feedback

[0313] Subject: Server

[0314] The server improves the generative AI model based on the feedback information received from viewers. The input is the feedback data, and the output is an improved generative AI model.

[0315] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0316] overview

[0317] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns them NFTs to preserve the value of the originals. Furthermore, by combining it with an emotion engine that recognizes user emotions, it is possible to analyze the impact of the generated voice on the viewer and improve the AI ​​model based on that feedback. This system aims to maintain the quality of anime works even when voice actors are unable to perform, thereby stabilizing their earnings.

[0318] Overall system configuration

[0319] The system of the present invention comprises the following elements:

[0320] 1. Voice analysis method: Analyze the voice actor's voice data.

[0321] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0322] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0323] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0324] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0325] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0326] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0327] 8. Emotion Engine: Recognizes and analyzes viewer emotions and uses that data to improve generative AI models.

[0328] Explanation of system processing

[0329] 1. Collection and analysis of audio data

[0330] Subject: Terminal

[0331] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[0332] The device uploads the recorded audio data to the server.

[0333] Subject: Server

[0334] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[0335] 2. Training a generative AI model

[0336] Subject: Server

[0337] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[0338] 3. Evaluation of generated speech

[0339] Subject: Server

[0340] The server generates new voices using generative AI and tests them to evaluate their quality.

[0341] The AI ​​model is repeatedly improved based on the evaluation results.

[0342] 4. NFT assignment to original audio

[0343] Subject: Server

[0344] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[0345] Registered NFTs prove ownership and authenticity of audio data.

[0346] 5. Audio Use and Revenue Sharing

[0347] Subject: Terminal (animation production company)

[0348] The device downloads the generated audio data and NFT from the server through an appropriate interface.

[0349] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[0350] Subject: Server

[0351] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payment.

[0352] 6. Audience Emotion Recognition and Feedback Collection

[0353] Subject: User (viewer)

[0354] Users can watch the completed animation and provide feedback to the platform on the quality of the generated voice and their impressions, and the emotion engine will recognize the user's emotions.

[0355] Subject: Server

[0356] The server receives and analyzes feedback from viewers and emotional data collected by the emotion engine, and the analysis results are used to further improve the AI ​​model.

[0357] Specific examples

[0358] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the voice actor's past recordings on the device is uploaded to a server and analyzed there. An AI model is trained using the analyzed feature data, and the new voice generated is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work, and a portion of the revenue is returned to the voice actor. Viewers watch the work, and emotions are recognized by an emotion engine, and further feedback is provided. The server improves the AI ​​model based on the viewer's emotional data and feedback, improving the quality of the next generated voice.

[0359] In this way, the present invention enables the production of higher quality anime works by incorporating a feedback system that takes into account the emotions of viewers while maintaining the quality of the voice actors' voices. It also establishes an effective revenue model for voice actors and provides a system that allows for the stable provision of works even in unforeseen circumstances.

[0360] The processing flow will be explained below.

[0361] Step 1:

[0362] Audio data collection

[0363] Subject: Terminal

[0364] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[0365] Step 2:

[0366] Uploading audio data

[0367] Subject: Terminal

[0368] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure security.

[0369] Step 3:

[0370] Analysis of audio data

[0371] Subject: Server

[0372] The server receives the uploaded audio data and analyzes its characteristics, such as waveform, spectrum, pitch, volume, and intonation, and stores the analysis results in a database.

[0373] Step 4:

[0374] Training generative AI models

[0375] Subject: Server

[0376] The server uses the analyzed voice feature data to train a generative AI model, at which point the model acquires the voice actor's unique vocal timbre.

[0377] Step 5:

[0378] Generating synthetic speech

[0379] Subject: Server

[0380] The server uses the trained generative AI model to generate new voice samples, which are then compared and evaluated against the original voice.

[0381] Step 6:

[0382] Evaluation of generated speech

[0383] Subject: Server

[0384] The server evaluates the quality of the generated speech, including similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[0385] Step 7:

[0386] NFT granting

[0387] Subject: Server

[0388] The server generates an NFT for the voice actor's original audio data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the audio data.

[0389] Step 8:

[0390] Download generated audio

[0391] Subject: Terminal (animation production company)

[0392] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[0393] Step 9:

[0394] Anime production and revenue records

[0395] Subject: Terminal

[0396] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[0397] Step 10:

[0398] Revenue sharing

[0399] Subject: Server

[0400] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process is carried out via electronic payment.

[0401] Step 11:

[0402] Emotion Recognition and Analysis

[0403] Subject: User (viewer)

[0404] The user watches an anime work, and the emotion engine recognizes the emotion they feel. The user's emotion data is then sent to the server.

[0405] Subject: Server

[0406] The server receives and analyzes the emotion data sent by the user, and the analysis results are used to improve the generative AI model.

[0407] Step 12:

[0408] Gathering feedback

[0409] Subject: User (viewer)

[0410] Users provide feedback to the platform about the quality and impressions of the generated speech, and the feedback is sent to the server.

[0411] Step 13:

[0412] Feedback analysis and model refinement

[0413] Subject: Server

[0414] The server receives and analyzes feedback from viewers and emotion analysis data from the emotion engine. The analysis results are used to improve the generative AI model and increase the quality of the next generated voice.

[0415] Examples:

[0416] Suppose the voice actor playing main character A in a long-running anime series has to take a break due to illness. Past voice data from the voice actor is recorded on a device and uploaded to a server. The server analyzes it and trains a generative AI model to generate new voice. The generated voice is then evaluated for quality and assigned an NFT. The animation production company downloads the generated voice and produces and releases the work, with a portion of the profits going back to the voice actor. Viewers watch the work, and the emotion engine analyzes their emotions, sending feedback to the server. The server uses the analysis results to improve the AI ​​model and increase the quality of the next generated voice. This system simultaneously maintains the quality of the voice actor's voice while also improving it to match the viewer's emotions.

[0417] Example 2

[0418] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0419] In traditional anime production, voice recording by voice actors is essential, and missing voice recording directly affects the quality of the work. If a voice actor is unable to participate in recording due to illness or other reasons, the risk of production delays increases. Voice actors' earnings also tend to be unstable. Furthermore, it is difficult to generate voices that reflect the viewer's emotions, making it difficult to consistently provide high-quality voices. Our goal is to solve these issues and provide a system that can stably provide higher-quality anime works.

[0420] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0421] In this invention, the server includes a voice analysis means, a training means for the AI ​​model, a means for evaluating the generated voice, a means for assigning NFTs to the voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion engine means for recognizing and analyzing viewers' emotional data, and a testing means for evaluating the quality of the generated voice. This allows for the generation of high-quality voice even when the voice actor is unable to participate in recording, and enables the generation of voice that reflects the viewer's emotions while maintaining the quality of the work. Furthermore, by steadily returning revenue to the voice actor, economic stability can be achieved.

[0422] A "voice analysis means" is a means for recording a voice actor's voice data and breaking down and analyzing the voice into characteristics such as waveform, spectrum, pitch, volume, and intonation.

[0423] "Means for training a generative AI model" refers to means for training a generative AI model using analyzed voice feature data so that it can reproduce the unique texture and tone of a voice actor.

[0424] The "means for evaluating generated speech" is a means for evaluating whether the generated speech is of appropriate quality, and uses speech quality evaluation metrics.

[0425] "Method of assigning NFTs to audio data" is a method of generating NFTs (non-fungible tokens) for a voice actor's original audio data, registering the NFTs on the blockchain, and proving the ownership and authenticity of the audio data.

[0426] "Means for downloading generated audio" refers to the means by which animation production companies can download the generated audio data and NFTs to their own devices.

[0427] A "revenue distribution vehicle" is a means for recording revenue earned after the release of an anime work and returning it to voice actors based on a preset percentage.

[0428] The "means for collecting feedback from viewers" refers to a means for collecting and analyzing feedback and emotional data provided by viewers after watching an anime work.

[0429] The "emotion engine means" is a means for recognizing and analyzing the viewer's emotions, and the data is used to improve the generative AI model.

[0430] "Testing means for evaluating the quality of generated speech" refers to a means for conducting tests to evaluate the quality of speech generated by a trained AI model.

[0431] MODE FOR CARRYING OUT THE INVENTION

[0432] This invention is a system that uses generative AI technology to replicate a voice actor's voice and assigns an NFT to preserve the value of the original. Furthermore, this system combines an emotion engine that recognizes the user's emotions, analyzes the impact of the generated voice on the viewer, and can improve the AI ​​model based on that feedback. This system can maintain the quality of anime works and stabilize the voice actor's income even when the voice actor is unable to participate in recording.

[0433] Overall system configuration

[0434] The system of the present invention comprises the following elements:

[0435] 1. Voice analysis method: Analyze the voice actor's voice data.

[0436] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0437] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0438] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0439] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0440] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0441] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0442] 8. Emotion engine means: Recognize and analyze viewer emotions and use that data to improve the generative AI model.

[0443] 9. Testing method to evaluate the quality of generated speech: Evaluate the quality of speech generated by the trained AI model.

[0444] Hardware and software used

[0445] Hardware

[0446] Recording device: A high-quality microphone (e.g., Shure SM7B), a PC for recording audio (e.g., a general-purpose laptop)

[0447] Server: A high-performance server in a data center (e.g., a virtual machine in a cloud environment)

[0448] Animation production device: PC for animation production (e.g., workstation equipped with an LCD pen tablet)

[0449] software

[0450] Voice analysis software: General-purpose voice analysis tools

[0451] Generative AI Models: A General-Purpose Deep Learning Framework

[0452] NFT Blockchain Platform: A Popular Distributed Ledger Technology

[0453] Electronic payment system: Online payment service

[0454] Feedback Collection Platform: Survey Tools

[0455] Explanation of system processing

[0456] The specific processing flow of this system will be explained below.

[0457] Audio data collection and analysis

[0458] Subject: Terminal

[0459] The device records the voice actor's voice using a high-quality microphone and generates an audio file.

[0460] The collected audio files contain the characters' lines and express various emotions.

[0461] The device uploads the audio file to the server.

[0462] Subject: Server

[0463] The server uses audio analysis software to analyze the received audio files.

[0464] The server breaks down the audio into features such as waveform, spectrum, pitch, volume, and intonation, and generates analysis data.

[0465] The server stores the analyzed data in a database.

[0466] Training generative AI models

[0467] Subject: Server

[0468] The server trains a generative AI model based on the analyzed voice feature data.

[0469] The generative AI model runs on a common deep learning framework.

[0470] Once trained, the generative AI model will be able to reproduce the unique texture and tone of a voice actor.

[0471] Evaluation and improvement of generated speech

[0472] Subject: Server

[0473] The server evaluates the quality of the generated voice to test new voices created by the generative AI model.

[0474] Voice quality assessment metrics are used in the testing.

[0475] Based on the test results, the parameters of the AI ​​model are readjusted and retrained.

[0476] NFT assignment for original audio

[0477] Subject: Server

[0478] The server generates an NFT for the voice actor's original voice.

[0479] The generated NFT is registered on the blockchain, proving ownership and authenticity of the audio data.

[0480] Audio Usage and Revenue Sharing

[0481] Subject: Terminal (animation production company)

[0482] The animation production company's terminal downloads the generated audio data and NFT from the server.

[0483] The downloaded audio data will be used in the production of anime works.

[0484] After the anime work is released, revenue data is uploaded from the terminal to the server.

[0485] Subject: Server

[0486] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[0487] An electronic payment system will be used for revenue sharing.

[0488] Audience emotion recognition and feedback collection

[0489] Subject: User (viewer)

[0490] Users can view the completed animation and enter their impressions in a feedback form.

[0491] The emotion engine recognizes and collects audience emotional data.

[0492] Subject: Server

[0493] The server analyzes the feedback from viewers and the emotional data from the emotion engine.

[0494] The analysis results will be used to further improve the AI ​​model, contributing to improving the quality of generated speech.

[0495] Examples and prompts

[0496] For example, this system would be effective if a voice actor playing a key character in an anime had to take a break due to illness. The voice data of the voice actor recorded on the device is uploaded to a server and analyzed. A generative AI model is trained based on the analysis results, and a new voice is generated. The original voice is assigned an NFT and registered on the blockchain. The animation production company downloads the generated voice and completes the work. Revenue is returned to the voice actor, and viewer emotional data and feedback are used to improve the AI ​​model.

[0497] Example prompt sentence:

[0498] "Generate the following line in the voice of the specified voice actor. Line: 'Hello, this is Character A. It's a great day today.'"

[0499] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0500] Divide the processing flow of this system's program into processing steps

[0501] Step 1: Collect and upload audio data

[0502] Step 2: Analyzing the audio data

[0503] Step 3: Training the generative AI model

[0504] Step 4: Evaluate and improve the generated speech

[0505] Step 5: Adding an NFT to the original audio

[0506] Step 6: Download and use the generated audio

[0507] Step 7: Revenue sharing

[0508] Step 8: Recognize audience emotions and collect feedback

[0509] Description of each processing step

[0510] Step 1: Collect and upload audio data

[0511] Subject: Terminal

[0512] The device uses a high-quality microphone to record voice data from voice actors, including the character's lines and various emotional expressions.

[0513] The device temporarily stores the recorded audio data and prepares it for uploading to the server.

[0514] Upload the audio file to the server using a secure communication protocol (e.g. HTTPS).

[0515] Input: Recorded audio file

[0516] Output: Audio file uploaded to the server

[0517] Step 2: Analyzing the audio data

[0518] Subject: Server

[0519] The server launches voice analysis software to analyze the received voice data.

[0520] The audio data is decomposed into features such as waveform, spectrum, pitch, volume, and intonation, and these data are generated.

[0521] The server stores the generated feature data in a database.

[0522] Input: Uploaded audio file

[0523] Output: Audio feature data

[0524] Step 3: Training the generative AI model

[0525] Subject: Server

[0526] The server uses the analyzed audio feature data to train a generative AI model.

[0527] Optimize the parameters of AI models using deep learning frameworks (e.g., TensorFlow, PyTorch).

[0528] The training process requires a lot of data and computational resources to reproduce the unique texture of the voice actor.

[0529] Input: Audio feature data

[0530] Output: A trained generative AI model

[0531] Step 4: Evaluate and improve the generated speech

[0532] Subject: Server

[0533] The server uses a generative AI model to generate new speech and evaluate its quality.

[0534] We use voice quality evaluation metrics (e.g., PESQ, STOI) and improve the AI ​​model based on the evaluation results.

[0535] The server repeats this process to improve the accuracy of the model.

[0536] Input: A trained generative AI model

[0537] Output: Evaluated generated speech data

[0538] Step 5: Adding an NFT to the original audio

[0539] Subject: Server

[0540] The server generates an NFT (non-fungible token) for the voice actor's original voice data.

[0541] NFTs are created and registered on a blockchain platform (e.g., Ethereum).

[0542] The generated NFT proves ownership and authenticity of the audio data.

[0543] Input: Original audio data

[0544] Output: NFT registered on the blockchain

[0545] Step 6: Download and use the generated audio

[0546] Subject: Terminal (animation production company)

[0547] The animation production company's terminal downloads the generated audio data and NFT from the server.

[0548] The device uses the downloaded audio data and applies it to the animation work being produced.

[0549] Input: Generated audio data, NFT

[0550] Output: Audio data and NFTs used as part of the anime production

[0551] Step 7: Revenue sharing

[0552] Subject: Terminal (animation production company)

[0553] The device records revenue data for released anime works.

[0554] Periodically upload revenue data to a server.

[0555] Subject: Server

[0556] The server receives the revenue data and distributes the revenue to the voice actors based on pre-set percentages.

[0557] Revenue sharing will be via electronic payment systems (e.g., online payment services).

[0558] Input: Revenue Data

[0559] Output: Revenue share to voice actors

[0560] Step 8: Recognize audience emotions and collect feedback

[0561] Subject: User (viewer)

[0562] Users can view the completed animation and enter their impressions in a feedback form.

[0563] The emotion engine recognizes and collects audience emotional data.

[0564] Input: Viewer sentiment data and feedback

[0565] Output: Emotion data and feedback collected on the server

[0566] Subject: Server

[0567] The server analyzes the feedback from viewers and the emotional data collected by the emotion engine.

[0568] The analysis results are used to improve the AI ​​model, leading to improved quality of the generated voice.

[0569] Input: Collected emotion data and feedback

[0570] Output: An improved generative AI model

[0571] (Application example 2)

[0572] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0573] In anime production, if a key voice actor is temporarily unavailable, it is difficult to complete the work while maintaining the quality of that character's voice. Furthermore, while viewer feedback is important for improving the quality of voice data generated by generative AI models, traditional feedback collection methods have the problem of being unable to recognize user emotions in detail. Furthermore, there is a need for a mechanism to guarantee the authenticity and ownership of generated voices and to appropriately distribute revenue.

[0574] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice analysis means, a means for training the generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion recognition means, and a means for analyzing the collected emotion data and improving the generative AI model based on the analysis results. This makes it possible to generate high-quality voice even when the main voice actor is unavailable, collect feedback based on viewer emotions, improve the generative AI model, and improve the quality of the generated voice. Furthermore, assigning NFTs to the generated voice guarantees the authenticity and ownership of the voice data and enables appropriate revenue distribution.

[0575] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[0576] "Means for training a generative AI model" means means for training a generative AI model based on analyzed audio data.

[0577] "Means for evaluating generated voice" refers to a means for evaluating the quality of voice generated by the generation AI.

[0578] "Means for assigning NFTs to audio data" refers to a means for assigning NFTs to generated audio data and registering it on the blockchain.

[0579] The "means for downloading generated voice" is a means for downloading generated voice data.

[0580] "Revenue distribution means" means a means for recording and managing revenues from generated voices and animation works, and distributing revenues to voice actors and related parties.

[0581] The "means for collecting feedback from viewers" is a means for collecting feedback from viewers and analyzing the data.

[0582] An "emotion recognition means" is a means for recognizing the viewer's emotions and collecting data on them.

[0583] "Means for analyzing collected emotional data and improving the generative AI model based on the results of the analysis" refers to means for analyzing emotional data collected from viewers and improving the generative AI model based on the results of the analysis.

[0584] The system of the present invention is realized by combining various means as follows. Specific embodiments are shown below.

[0585] 1. Collection and analysis of audio data

[0586] Subject: Terminal

[0587] Voice actors' voice data is recorded using high-quality recording equipment, and this voice includes the characters' lines and various emotional expressions.

[0588] The recorded audio data is uploaded from the device to a server, using high security and data compression technology to prevent data loss.

[0589] Subject: Server

[0590] The server receives the uploaded audio data and analyzes characteristics such as the audio waveform, spectrum, pitch, and tempo. This analysis is performed using libraries such as "librosa."

[0591] The analyzed audio feature data is stored in a database for use in training subsequent generative AI models.

[0592] 2. Training a generative AI model

[0593] Subject: Server

[0594] The server trains a generative AI model based on the speech feature data for analysis, using machine learning frameworks such as the "transformers" library.

[0595] The model is tuned to reproduce the unique texture of a particular voice actor's voice.

[0596] 3. Evaluation of generated speech

[0597] Subject: Server

[0598] The server generates new audio using the trained generative AI model and evaluates its quality, applying criteria for voice quality assessment and further fine-tuning the model if necessary.

[0599] 4. NFT assignment to original audio

[0600] Subject: Server

[0601] The server assigns an NFT to the voice actor's original voice data and registers the NFT on the blockchain, using an external blockchain API.

[0602] A registered NFT serves as proof of ownership and authenticity of the audio data.

[0603] 5. Audio Use and Revenue Sharing

[0604] Subject: Terminal (animation production company)

[0605] The animation production company downloads the generated audio data and NFTs from the server using a dedicated interface.

[0606] The downloaded generated audio is used to create an animated work, and after the work is released, revenue data is recorded and sent to a server.

[0607] Subject: Server

[0608] The server returns the revenue to the voice actor based on the received revenue data at a preset rate, using electronic payment technology.

[0609] 6. Audience Emotion Recognition and Feedback Collection

[0610] Subject: User (viewer)

[0611] Users can watch the finished animation and provide feedback on the quality of the generated voice, while the system recognizes the user's emotions in real time using the "transformers" library.

[0612] Subject: Server

[0613] The server collects and analyzes viewer feedback and emotion recognition results, which are used to improve the next generation AI model.

[0614] Specific examples

[0615] For example, there may be an anime featuring a character played by a well-known voice actor, but that voice actor is suddenly unable to appear. In this case, previously recorded voice data expressing the actor's lines and emotions is uploaded to a server, and a generative AI model is trained based on that data. Furthermore, by entering a prompt sentence as follows, character voices for specific situations can be generated.

[0616] (Example prompt): "Good morning, let's do our best today!" (in a cheerful tone)

[0617] This voice is used in the animation, and viewers can provide emotional feedback, such as "I'm excited." This feedback is analyzed in detail using emotion recognition technology and used to further improve the quality of the generative AI model.

[0618] In this way, the system of the present invention can provide high-quality generated audio that takes into account the emotions of the viewer, allowing for smooth animation production.

[0619] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0620] Step 1:

[0621] Subject: Terminal

[0622] High-quality recording equipment is used to record voice data from voice actors, including the characters' lines and various emotional expressions.

[0623] The recorded audio data is uploaded from the device to the server, using high security and data compression technology to prevent data loss.

[0624] Input: Voice actor's voice data

[0625] Output: Uploaded audio data

[0626] Step 2:

[0627] Subject: Server

[0628] The server receives the uploaded audio data.

[0629] The received audio data is analyzed and features such as audio waveform, spectrum, pitch, and tempo are extracted.

[0630] The extracted feature data is stored in a database. The "librosa" library is used to extract audio features.

[0631] Input: Uploaded audio data

[0632] Output: Analyzed audio feature data

[0633] Step 3:

[0634] Subject: Server

[0635] A generative AI model is trained based on the analyzed audio feature data using a machine learning framework such as the "transformers" library.

[0636] The trained generative AI model is tuned to reproduce the unique texture of a particular voice actor's voice.

[0637] Input: Analyzed speech feature data

[0638] Output: A trained generative AI model

[0639] Step 4:

[0640] Subject: Server

[0641] It uses trained generative AI models to generate new voices, using prompts to tailor voices to specific tones and situations.

[0642] Example prompt: "Good morning, let's do our best today!" (in a cheerful tone)

[0643] Input: Trained generative AI model, prompt

[0644] Output: Generated audio data

[0645] Step 5:

[0646] Subject: Server

[0647] The quality of the generated speech is evaluated, applying criteria for speech quality assessment and further fine-tuning the model if necessary.

[0648] Input: Generated audio data

[0649] Output: Evaluation results, improved generative AI model

[0650] Step 6:

[0651] Subject: Server

[0652] The generated audio data is assigned an NFT and registered on the blockchain using an external blockchain API.

[0653] Input: Generated audio data

[0654] Output: Audio data with NFT attached

[0655] Step 7:

[0656] Subject: Terminal (animation production company)

[0657] Download the generated audio data and NFT from the server using a dedicated interface.

[0658] Create an animated work using the downloaded generated audio.

[0659] Input: Audio data with NFT attached

[0660] Output: Anime work

[0661] Step 8:

[0662] Subject: Terminal (animation production company)

[0663] After the anime work is released, revenue data is recorded and sent to a server.

[0664] Input: Revenue Data

[0665] Output: Revenue data sent to the server

[0666] Step 9:

[0667] Subject: Server

[0668] Based on the received revenue data, revenue is returned to the voice actor based on a preset percentage, and electronic payment technology is used here.

[0669] Input: Revenue Data

[0670] Output: Revenue share to voice actors

[0671] Step 10:

[0672] Subject: User (viewer)

[0673] View the finished animation and provide feedback on the quality of the generated audio.

[0674] The emotions of the user are recognized in real time while watching. The "transformers" library is used for emotion recognition.

[0675] Input: Viewer feedback, emotion recognition data

[0676] Output: Feedback data, emotion recognition data

[0677] Step 11:

[0678] Subject: Server

[0679] Collect and analyze viewer feedback and emotion recognition results.

[0680] The analysis results will be used to improve the next generative AI model.

[0681] Input: Feedback data, emotion recognition data

[0682] Output: Analysis results, improved generative AI model

[0683] The above are the specific processing steps for carrying out the present invention. The specific operations performed in each step, the software used, and the data processing process have been shown.

[0684] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0685] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0686] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0687] [Second embodiment]

[0688] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0689] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0690] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0691] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0692] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0693] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0694] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0695] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0696] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0697] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0698] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0699] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0700] overview

[0701] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns NFTs to them to preserve the value of the originals. This aims to maintain the quality of anime works even when voice actors are unable to perform, and to stabilize voice actors' earnings.

[0702] Overall system configuration

[0703] The system of the present invention comprises the following elements:

[0704] 1. Voice analysis method: Analyze the voice actor's voice data.

[0705] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0706] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0707] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0708] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0709] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0710] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0711] Explanation of system processing

[0712] 1. Collection and analysis of audio data

[0713] Subject: Terminal

[0714] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[0715] The device uploads the recorded audio data to the server.

[0716] Subject: Server

[0717] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[0718] 2. Training a generative AI model

[0719] Subject: Server

[0720] The server uses the analyzed audio feature data to train a generative AI model.

[0721] This AI model is designed to reproduce the unique vocal qualities of voice actors.

[0722] 3. Evaluation of generated speech

[0723] Subject: Server

[0724] The server generates new voices using generative AI and tests them to evaluate their quality.

[0725] The model is repeatedly improved based on the evaluation results.

[0726] 4. NFT assignment to original audio

[0727] Subject: Server

[0728] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[0729] Registered NFTs prove ownership and authenticity of audio data.

[0730] 5. Audio Use and Revenue Sharing

[0731] Subject: Terminal (animation production company)

[0732] The device downloads the generated audio data and NFT from the server.

[0733] Use synthetic audio in animation production and release the finished work.

[0734] The revenue of the work is recorded and sent to the server.

[0735] Subject: Server

[0736] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[0737] 6. Quality check and feedback

[0738] Subject: User (viewer)

[0739] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[0740] Subject: Server

[0741] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[0742] Specific examples

[0743] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This system eliminates the decline in quality and instability of revenue that can occur when voice actors are replaced.

[0744] The processing flow will be explained below.

[0745] Step 1:

[0746] Audio data collection

[0747] Subject: Terminal

[0748] The device records voice actors' voices in high quality, including the characters' lines and various emotional expressions.

[0749] Step 2:

[0750] Uploading audio data

[0751] Subject: Terminal

[0752] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure secure transmission.

[0753] Step 3:

[0754] Analysis of audio data

[0755] Subject: Server

[0756] The server analyzes the received audio data. This analysis involves extracting features such as the waveform, spectrum, pitch, volume, and intonation of the audio. The analysis results are stored in a database.

[0757] Step 4:

[0758] Training generative AI models

[0759] Subject: Server

[0760] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[0761] Step 5:

[0762] Generating synthetic speech

[0763] Subject: Server

[0764] The server uses a trained generative AI model to generate new voice samples, which are then compared to the original voice.

[0765] Step 6:

[0766] Evaluation of generated speech

[0767] Subject: Server

[0768] The server evaluates the quality of the generated voice, including voice similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[0769] Step 7:

[0770] NFT granting

[0771] Subject: Server

[0772] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the voice data.

[0773] Step 8:

[0774] Download generated audio

[0775] Subject: Terminal (animation production company)

[0776] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[0777] Step 9:

[0778] Anime production and revenue records

[0779] Subject: Terminal

[0780] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[0781] Step 10:

[0782] Revenue sharing

[0783] Subject: Server

[0784] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payments.

[0785] Step 11:

[0786] Gathering feedback from viewers

[0787] Subject: User

[0788] Users can watch the completed animation and provide feedback to the platform on the quality of the generated audio and their impressions.

[0789] Step 12:

[0790] Feedback analysis and AI model improvement

[0791] Subject: Server

[0792] The server collects and analyzes feedback from viewers, and the analysis results are used to further refine the AI ​​model, thereby improving the quality of the generated voice in future iterations.

[0793] Example 1

[0794] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0795] If a voice actor is unable to provide voice over for illness or other reasons, the quality of the anime work declines and revenue becomes unstable. Furthermore, there is no way to guarantee the authenticity of the original voice over, which creates the risk of counterfeiting or unauthorized use. Therefore, a method is needed to effectively collect viewer feedback and improve generative AI models.

[0796] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0797] In this invention, the server includes a voice analysis means for analyzing voice data and extracting feature data; a generative AI model training means for training a generative AI model based on the analyzed feature data; a generated voice evaluation means for generating new voices using the generative AI model and evaluating their quality; a voice data NFT assignment means for assigning NFTs to the voice actor's original voice data and registering it on the blockchain; a revenue distribution means for returning revenue to the voice actor based on revenue data; and a feedback collection means for collecting feedback from viewers and using it to improve the generative AI model. This enables the production of high-quality anime works even when voice actors are unavailable, improving revenue stability. It also ensures the authenticity of the original voice and enables continuous improvement of the AI ​​model through feedback.

[0798] 1. "Audio data" means data that is a digital recording of the sounds and lines spoken by a voice actor.

[0799] 2. "Terminal means" means a hardware device or software tool for recording or uploading audio data.

[0800] 3. "Server" refers to a computer system that performs processes such as analyzing voice data, evaluating generated voice data, and assigning NFTs.

[0801] 4. "Audio analysis means" means technologies or algorithms that process uploaded audio data to extract characteristics such as waveform, spectrum, pitch, volume, and intonation.

[0802] 5. A "generative AI model" is a machine learning model that generates new speech based on speech data.

[0803] 6. "Generative AI model training means" means the process of optimizing and training a generative AI model using analyzed speech feature data.

[0804] 7. “Generative Speech Evaluation Measures” means tests or criteria for assessing the quality of speech generated by a generative AI model.

[0805] 8. "NFT" stands for "Non-Fungible Token" and is a token used to prove ownership and authenticity of digital assets.

[0806] 9. "Audio data NFT assignment method" is a technology that generates NFTs for audio data and registers them on the blockchain.

[0807] 10. "Revenue sharing mechanism" is a system for recording revenue from anime works and distributing that revenue fairly to voice actors.

[0808] 11. “Feedback collection methods” are technologies and methods used to collect audience ratings and feedback and use it to improve the generative AI model.

[0809] MODE FOR CARRYING OUT THE INVENTION

[0810] This invention is a system that uses generative AI technology to replicate the voice of a voice actor and assigns an NFT to preserve the value of the original. Specific embodiments for implementing this system are described below.

[0811] 1. Hardware Configuration

[0812] Terminal

[0813] Use high-quality dedicated microphones and studio equipment to record audio data, and a computer or dedicated device to capture the recorded audio data in digital format (e.g., WAV files) and upload it to a server via an internet connection.

[0814] server

[0815] It uses a high-performance server that analyzes voice data, trains generative AI models, evaluates generated voices, assigns NFTs, distributes revenue, and collects feedback. Specifically, it refers to a server with the computing resources to run libraries such as Python, TensorFlow, and PyTorch.

[0816] 2. Software Configuration

[0817] Voice analysis methods

[0818] We use LibROSA or other audio analysis libraries to extract features from the audio data, such as waveform, spectrum, pitch, volume, and intonation, and store these in a database.

[0819] Generative AI model training tools

[0820] To train the generative AI model, we use Python and machine learning libraries such as TensorFlow and PyTorch. Using the analyzed audio feature data, we optimize the model (e.g., WaveNet, Tacotron2) to reproduce the unique voice quality of the voice actor.

[0821] Generated speech evaluation means

[0822] To evaluate the quality of the generated speech, we evaluate the performance of the model using speech evaluation criteria such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[0823] Audio data NFT granting method

[0824] To assign an NFT to audio data, a smart contract creation tool (e.g., Solidity) is used and registered on the Ethereum blockchain to prove the authenticity and ownership of the audio data.

[0825] Revenue sharing method

[0826] Revenue data is collected through online payment systems (e.g., PayPal, Stripe), and revenue is distributed to voice actors based on that data. Revenue distribution is calculated using automated scripts and programs on the server.

[0827] Feedback collection methods

[0828] A dedicated feedback form and application are used to collect feedback from viewers, and the collected data is used as an evaluation index for the generative AI model, helping to improve the model.

[0829] Specific examples

[0830] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. The original voice is assigned an NFT to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This process prevents a decline in quality due to the voice actor's absence.

[0831] Prompt Sentence Examples

[0832] "As a first step, please record voice actor A's past voice data and upload it to our server. Next, we will use the server to analyze this voice data and train a generative AI model. Once the model is trained, we will evaluate the quality of the generated voice and make improvements if necessary. Finally, please attach an NFT to the original voice and download the generated voice for use."

[0833] This invention makes it possible to produce high-quality anime even when voice actors are unavailable, improving revenue stability and enabling continuous improvement of AI models based on viewer feedback.

[0834] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0835] Step 1: Record and upload audio data

[0836] Subject: Terminal

[0837] The device uses dedicated microphones and studio equipment to record high-quality voice data from voice actors. The recorded voice data (e.g., in WAV format) includes the character's lines and emotional expressions.

[0838] Once recording is complete, the device uploads the audio data to a server over the internet using a secure file transfer protocol (e.g., SFTP).

[0839] Input: Voice actor's voice data

[0840] Output: Upload audio data to the server

[0841] Step 2: Analyzing the audio data

[0842] Subject: Server

[0843] The server receives the uploaded audio data and uses an audio analysis library such as LibROSA to extract features such as waveform, spectrum, pitch, volume, and intonation.

[0844] The analyzed voice feature data is stored in a database.

[0845] Input: Uploaded audio data

[0846] Output: Analyzed audio feature data

[0847] Step 3: Training the generative AI model

[0848] Subject: Server

[0849] The server trains a generative AI model based on the analyzed voice feature data, using Python and machine learning libraries such as TensorFlow and PyTorch. The generative AI model employs a multilayer neural network-based model such as WaveNet or Tacotron2.

[0850] The model is trained to replicate the voice actor's unique vocal timbre and speaking style.

[0851] Input: Analyzed audio feature data

[0852] Output: A trained generative AI model

[0853] Step 4: Evaluate and improve the generated speech

[0854] Subject: Server

[0855] The server generates test audio using a trained generative AI model, which is then evaluated using audio metrics such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[0856] Based on the evaluation results, improvements are made repeatedly by adjusting the model's hyperparameters and retraining with additional data.

[0857] Input: Trained generative AI model, test audio data

[0858] Output: Evaluation results, improved generative AI model

[0859] Step 5: Adding an NFT to the original audio

[0860] Subject: Server

[0861] The server generates an NFT for the original audio data, using a smart contract to prove ownership and authenticity of the audio data.

[0862] The generated NFT is registered on the Ethereum blockchain and acts as a digital certificate.

[0863] Input: Original audio data

[0864] Output: NFT, registered on the blockchain

[0865] Step 6: Download and use the audio data

[0866] Subject: Terminal (animation production company)

[0867] The device downloads the generated audio data and NFT from the server, which is done securely using SFTP.

[0868] The animation production company will use the downloaded generated audio to create and release an animated work.

[0869] Input: Generated audio data, NFT

[0870] Output: Finished animation

[0871] Step 7: Revenue sharing

[0872] Subject: Terminal (animation production company)

[0873] The terminal records the revenue generated by the published anime works and transmits the revenue data to the server, which collects the data through online payment systems (e.g., PayPal, Stripe).

[0874] Input: Revenue data for completed anime works

[0875] Output: Send revenue data to the server

[0876] Subject: Server

[0877] The server distributes revenue to voice actors based on the received revenue data. Revenue distribution calculations are performed using automated scripts or programs based on pre-set percentages.

[0878] Input: Revenue data for anime works

[0879] Output: Revenue return to voice actors

[0880] Step 8: Gather viewer feedback and refine the model

[0881] Subject: User (viewer)

[0882] After watching the anime, users can provide their impressions and audio quality ratings using a dedicated feedback form or app.

[0883] Input: Viewer feedback

[0884] Output: Feedback data

[0885] Subject: Server

[0886] The server collects feedback from viewers and analyzes it as an evaluation index for machine learning. Based on the results of this analysis, the generative AI model is further improved.

[0887] Input: Feedback data

[0888] Output: An improved generative AI model

[0889] (Application example 1)

[0890] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0891] Conventional voice generation systems for voice actors have had problems maintaining the quality of voice data and stable revenue when voice actors are unable to physically participate in recording. Furthermore, there is no way to guarantee the authenticity or ownership of the generated voice data, making it difficult to effectively collect feedback from viewers and use it to improve AI models. The present invention aims to solve these problems by stably generating voice data for voice actors, ensuring revenue, and effectively utilizing feedback from viewers.

[0892] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0893] In this invention, the server includes a voice analysis means, a training means for a generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, a means for listening to the voice data through a listening application installed on a smartphone, a means for purchasing NFTs via the listening application, and a means for collecting feedback information via the listening application. This allows the quality of the generated voice and stability of revenue to be maintained even in situations where the voice actor cannot participate in recording, and enables the AI ​​model to be improved based on feedback from viewers.

[0894] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[0895] A "training method for a generative AI model" is a method for training an AI model based on analyzed voice data.

[0896] The "means for evaluating generated speech" is a means for evaluating and testing the quality of generated speech.

[0897] "Method of assigning NFTs to audio data" is a method of assigning NFTs to voice actor audio data to prove its authenticity and ownership.

[0898] The "means for downloading generated voice" is a means for making the generated voice data downloadable.

[0899] The "profit distribution means" is a means for appropriately distributing the profits from the content in which the generated audio is used.

[0900] "Means for collecting feedback from viewers" refers to means for collecting opinions and impressions from viewers.

[0901] "Means for listening to audio data through a listening application installed on a smartphone" refers to means for listening to audio data through an application installed on a smartphone.

[0902] "Means for purchasing NFTs via a viewing application" refers to means for purchasing NFTs using a viewing application.

[0903] The "means for collecting feedback information via a viewing application" refers to a means for collecting feedback information from viewers through a viewing application.

[0904] System configuration

[0905] The system of the present invention comprises the following elements:

[0906] 1. Voice analysis method: Analyze the voice actor's voice data.

[0907] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0908] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0909] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0910] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0911] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0912] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0913] 8. Means for listening to audio data through a listening application installed on a smartphone: A smartphone application is used as a means for a user to listen to audio data.

[0914] 9. Means for purchasing NFTs via the viewing application: A means for users to purchase NFTs via the viewing application.

[0915] 10. Means of collecting feedback information via a viewing application: Means of collecting viewer feedback information through a viewing application.

[0916] Explanation of program processing

[0917] 1. Collection and analysis of audio data

[0918] Subject: Terminal

[0919] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[0920] The device uploads the recorded audio data to the server.

[0921] Subject: Server

[0922] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and modulation, and stores the data in a database.

[0923] 2. Training a generative AI model

[0924] Subject: Server

[0925] The server uses the analyzed audio feature data to train a generative AI model, which is built using libraries such as TensorFlow and PyTorch.

[0926] 3. Evaluation of generated speech

[0927] Subject: Server

[0928] The server generates new voices using the generative AI and performs tests to evaluate their quality. The model is then repeatedly improved based on the evaluation results. A Python evaluation algorithm is used for the evaluation.

[0929] 4. NFT assignment to original audio

[0930] Subject: Server

[0931] The server generates an NFT using Ethereum for the voice actor's original voice data and registers it on the blockchain via the OpenSea API.

[0932] 5. Audio Use and Revenue Sharing

[0933] Subject: Terminal (animation production company)

[0934] The device downloads the generated audio data and NFTs from the server, uses the generated audio in animation production, and publishes the finished work. The revenue from the work is recorded and sent to the server.

[0935] Subject: Server

[0936] The server receives the revenue data and distributes the revenue to the voice actors based on a preset percentage. AWS is used as the revenue management system.

[0937] 6. Quality check and feedback

[0938] Subject: User (viewer)

[0939] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[0940] Subject: Server

[0941] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[0942] Examples of specific examples and prompts

[0943] For example, we will show a specific example using the smartphone app "Spoken NFT."

[0944] 1. App launch: The user launches the "Spoken NFT" app and searches for the generated voice of their favorite voice actor.

[0945] 2. Listen to the audio: Select the generated audio from the search results and listen to it on your smartphone.

[0946] 3. Purchase NFT: Purchase an NFT for the generated audio you like and retain ownership on the blockchain.

[0947] 4. Provide feedback: After listening, send feedback about the generated audio through the app.

[0948] Here are some example prompts to input to a generative AI model:

[0949] "Generate lines for character A with the characteristics of voice actor XX:

[0950] "From today onwards, you are a part of our team!"

[0951] Emotion: Joy

[0952] Based on this prompt, the AI ​​will reproduce the voice quality of voice actor XX and the characteristics of character A, and generate voice that also takes emotional expression into account.

[0953] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0954] Step 1: Collecting audio data

[0955] Subject: Terminal

[0956] The device records high-quality voice data from the voice actor. This voice includes the character's lines and emotional expressions. The device then uploads this voice data to the server. The input is the recorded voice data, and the output is the audio file uploaded to the server.

[0957] Step 2: Analyzing the audio data

[0958] Subject: Server

[0959] The server analyzes the received audio data by breaking it down into features such as waveform, spectrum, pitch, volume, and modulation. The analysis results are stored in a database. The input is the uploaded audio file, and the output is the analyzed audio feature data.

[0960] Step 3: Training the generative AI model

[0961] Subject: Server

[0962] The server uses the analyzed voice feature data to train a generative AI model using TensorFlow or PyTorch. During the training process, data processing and calculations are performed to update the AI ​​model. The input is the analyzed voice feature data, and the output is a trained generative AI model.

[0963] Step 4: Generate and evaluate synthetic speech

[0964] Subject: Server

[0965] The server generates new speech data using a trained generative AI model, then evaluates the quality of the generated speech and refines the AI ​​model as needed. The input is the trained generative AI model and a text prompt, and the output is the generated new speech data.

[0966] Step 5: Adding an NFT to the audio data

[0967] Subject: Server

[0968] The server generates an NFT for the generated audio data using Ethereum and registers it on the blockchain via the OpenSea API. The input is the generated audio data, and the output is the audio data with the NFT attached and a record of its ownership.

[0969] Step 6: Download the generated audio

[0970] Subject: Terminal (animation production company)

[0971] The device downloads the generated audio data and NFT from the server. The input is the audio data with the NFT attached, and the output is the downloaded audio file.

[0972] Step 7: Revenue sharing

[0973] Subject: Server

[0974] The server receives revenue data for the work and distributes the revenue to the voice actors based on a preset percentage. The input is revenue data, and the output is distributed revenue information.

[0975] Step 8: Gather feedback from your audience

[0976] Subject: User (viewer)

[0977] After listening to the generated speech through a smartphone application, the user sends feedback. The input is the feedback content, and the output is the feedback data sent to the server.

[0978] Step 9: Improve the AI ​​model with feedback

[0979] Subject: Server

[0980] The server improves the generative AI model based on the feedback information received from viewers. The input is the feedback data, and the output is an improved generative AI model.

[0981] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0982] overview

[0983] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns them NFTs to preserve the value of the originals. Furthermore, by combining it with an emotion engine that recognizes user emotions, it is possible to analyze the impact of the generated voice on the viewer and improve the AI ​​model based on that feedback. This system aims to maintain the quality of anime works even when voice actors are unable to perform, thereby stabilizing their earnings.

[0984] Overall system configuration

[0985] The system of the present invention comprises the following elements:

[0986] 1. Voice analysis method: Analyze the voice actor's voice data.

[0987] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[0988] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[0989] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[0990] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[0991] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[0992] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[0993] 8. Emotion Engine: Recognizes and analyzes viewer emotions and uses that data to improve generative AI models.

[0994] Explanation of system processing

[0995] 1. Collection and analysis of audio data

[0996] Subject: Terminal

[0997] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[0998] The device uploads the recorded audio data to the server.

[0999] Subject: Server

[1000] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[1001] 2. Training a generative AI model

[1002] Subject: Server

[1003] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[1004] 3. Evaluation of generated speech

[1005] Subject: Server

[1006] The server generates new voices using generative AI and tests them to evaluate their quality.

[1007] The AI ​​model is repeatedly improved based on the evaluation results.

[1008] 4. NFT assignment to original audio

[1009] Subject: Server

[1010] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[1011] Registered NFTs prove ownership and authenticity of audio data.

[1012] 5. Audio Use and Revenue Sharing

[1013] Subject: Terminal (animation production company)

[1014] The device downloads the generated audio data and NFT from the server through an appropriate interface.

[1015] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[1016] Subject: Server

[1017] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payment.

[1018] 6. Audience Emotion Recognition and Feedback Collection

[1019] Subject: User (viewer)

[1020] Users can watch the completed animation and provide feedback to the platform on the quality of the generated voice and their impressions, and the emotion engine will recognize the user's emotions.

[1021] Subject: Server

[1022] The server receives and analyzes feedback from viewers and emotional data collected by the emotion engine, and the analysis results are used to further improve the AI ​​model.

[1023] Specific examples

[1024] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the voice actor's past recordings on the device is uploaded to a server and analyzed there. An AI model is trained using the analyzed feature data, and the new voice generated is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work, and a portion of the revenue is returned to the voice actor. Viewers watch the work, and emotions are recognized by an emotion engine, and further feedback is provided. The server improves the AI ​​model based on the viewer's emotional data and feedback, improving the quality of the next generated voice.

[1025] In this way, the present invention enables the production of higher quality anime works by incorporating a feedback system that takes into account the emotions of viewers while maintaining the quality of the voice actors' voices. It also establishes an effective revenue model for voice actors and provides a system that allows for the stable provision of works even in unforeseen circumstances.

[1026] The processing flow will be explained below.

[1027] Step 1:

[1028] Audio data collection

[1029] Subject: Terminal

[1030] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[1031] Step 2:

[1032] Uploading audio data

[1033] Subject: Terminal

[1034] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure security.

[1035] Step 3:

[1036] Analysis of audio data

[1037] Subject: Server

[1038] The server receives the uploaded audio data and analyzes its characteristics, such as waveform, spectrum, pitch, volume, and intonation, and stores the analysis results in a database.

[1039] Step 4:

[1040] Training generative AI models

[1041] Subject: Server

[1042] The server uses the analyzed voice feature data to train a generative AI model, at which point the model acquires the voice actor's unique vocal timbre.

[1043] Step 5:

[1044] Generating synthetic speech

[1045] Subject: Server

[1046] The server uses the trained generative AI model to generate new voice samples, which are then compared and evaluated against the original voice.

[1047] Step 6:

[1048] Evaluation of generated speech

[1049] Subject: Server

[1050] The server evaluates the quality of the generated speech, including similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[1051] Step 7:

[1052] NFT granting

[1053] Subject: Server

[1054] The server generates an NFT for the voice actor's original audio data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the audio data.

[1055] Step 8:

[1056] Download generated audio

[1057] Subject: Terminal (animation production company)

[1058] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[1059] Step 9:

[1060] Anime production and revenue records

[1061] Subject: Terminal

[1062] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[1063] Step 10:

[1064] Revenue sharing

[1065] Subject: Server

[1066] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process is carried out via electronic payment.

[1067] Step 11:

[1068] Emotion Recognition and Analysis

[1069] Subject: User (viewer)

[1070] The user watches an anime work, and the emotion engine recognizes the emotion they feel. The user's emotion data is then sent to the server.

[1071] Subject: Server

[1072] The server receives and analyzes the emotion data sent by the user, and the analysis results are used to improve the generative AI model.

[1073] Step 12:

[1074] Gathering feedback

[1075] Subject: User (viewer)

[1076] Users provide feedback to the platform about the quality and impressions of the generated speech, and the feedback is sent to the server.

[1077] Step 13:

[1078] Feedback analysis and model refinement

[1079] Subject: Server

[1080] The server receives and analyzes feedback from viewers and emotion analysis data from the emotion engine. The analysis results are used to improve the generative AI model and increase the quality of the next generated voice.

[1081] Examples:

[1082] Suppose the voice actor playing main character A in a long-running anime series has to take a break due to illness. Past voice data from the voice actor is recorded on a device and uploaded to a server. The server analyzes it and trains a generative AI model to generate new voice. The generated voice is then evaluated for quality and assigned an NFT. The animation production company downloads the generated voice and produces and releases the work, with a portion of the profits going back to the voice actor. Viewers watch the work, and the emotion engine analyzes their emotions, sending feedback to the server. The server uses the analysis results to improve the AI ​​model and increase the quality of the next generated voice. This system simultaneously maintains the quality of the voice actor's voice while also improving it to match the viewer's emotions.

[1083] Example 2

[1084] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1085] In traditional anime production, voice recording by voice actors is essential, and missing voice recording directly affects the quality of the work. If a voice actor is unable to participate in recording due to illness or other reasons, the risk of production delays increases. Voice actors' earnings also tend to be unstable. Furthermore, it is difficult to generate voices that reflect the viewer's emotions, making it difficult to consistently provide high-quality voices. Our goal is to solve these issues and provide a system that can stably provide higher-quality anime works.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1087] In this invention, the server includes a voice analysis means, a training means for the AI ​​model, a means for evaluating the generated voice, a means for assigning NFTs to the voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion engine means for recognizing and analyzing viewers' emotional data, and a testing means for evaluating the quality of the generated voice. This allows for the generation of high-quality voice even when the voice actor is unable to participate in recording, and enables the generation of voice that reflects the viewer's emotions while maintaining the quality of the work. Furthermore, by steadily returning revenue to the voice actor, economic stability can be achieved.

[1088] A "voice analysis means" is a means for recording a voice actor's voice data and breaking down and analyzing the voice into characteristics such as waveform, spectrum, pitch, volume, and intonation.

[1089] "Means for training a generative AI model" refers to means for training a generative AI model using analyzed voice feature data so that it can reproduce the unique texture and tone of a voice actor.

[1090] The "means for evaluating generated speech" is a means for evaluating whether the generated speech is of appropriate quality, and uses speech quality evaluation metrics.

[1091] "Method of assigning NFTs to audio data" is a method of generating NFTs (non-fungible tokens) for a voice actor's original audio data, registering the NFTs on the blockchain, and proving the ownership and authenticity of the audio data.

[1092] "Means for downloading generated audio" refers to the means by which animation production companies can download the generated audio data and NFTs to their own devices.

[1093] A "revenue distribution vehicle" is a means for recording revenue earned after the release of an anime work and returning it to voice actors based on a preset percentage.

[1094] The "means for collecting feedback from viewers" refers to a means for collecting and analyzing feedback and emotional data provided by viewers after watching an anime work.

[1095] The "emotion engine means" is a means for recognizing and analyzing the viewer's emotions, and the data is used to improve the generative AI model.

[1096] "Testing means for evaluating the quality of generated speech" refers to a means for conducting tests to evaluate the quality of speech generated by a trained AI model.

[1097] MODE FOR CARRYING OUT THE INVENTION

[1098] This invention is a system that uses generative AI technology to replicate a voice actor's voice and assigns an NFT to preserve the value of the original. Furthermore, this system combines an emotion engine that recognizes the user's emotions, analyzes the impact of the generated voice on the viewer, and can improve the AI ​​model based on that feedback. This system can maintain the quality of anime works and stabilize the voice actor's income even when the voice actor is unable to participate in recording.

[1099] Overall system configuration

[1100] The system of the present invention comprises the following elements:

[1101] 1. Voice analysis method: Analyze the voice actor's voice data.

[1102] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[1103] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[1104] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[1105] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[1106] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[1107] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[1108] 8. Emotion engine means: Recognize and analyze viewer emotions and use that data to improve the generative AI model.

[1109] 9. Testing method to evaluate the quality of generated speech: Evaluate the quality of speech generated by the trained AI model.

[1110] Hardware and software used

[1111] Hardware

[1112] Recording device: A high-quality microphone (e.g., Shure SM7B), a PC for recording audio (e.g., a general-purpose laptop)

[1113] Server: A high-performance server in a data center (e.g., a virtual machine in a cloud environment)

[1114] Animation production device: PC for animation production (e.g., workstation equipped with an LCD pen tablet)

[1115] software

[1116] Voice analysis software: General-purpose voice analysis tools

[1117] Generative AI Models: A General-Purpose Deep Learning Framework

[1118] NFT Blockchain Platform: A Popular Distributed Ledger Technology

[1119] Electronic payment system: Online payment service

[1120] Feedback Collection Platform: Survey Tools

[1121] Explanation of system processing

[1122] The specific processing flow of this system will be explained below.

[1123] Audio data collection and analysis

[1124] Subject: Terminal

[1125] The device records the voice actor's voice using a high-quality microphone and generates an audio file.

[1126] The collected audio files contain the characters' lines and express various emotions.

[1127] The device uploads the audio file to the server.

[1128] Subject: Server

[1129] The server uses audio analysis software to analyze the received audio files.

[1130] The server breaks down the audio into features such as waveform, spectrum, pitch, volume, and intonation, and generates analysis data.

[1131] The server stores the analyzed data in a database.

[1132] Training generative AI models

[1133] Subject: Server

[1134] The server trains a generative AI model based on the analyzed voice feature data.

[1135] The generative AI model runs on a common deep learning framework.

[1136] Once trained, the generative AI model will be able to reproduce the unique texture and tone of a voice actor.

[1137] Evaluation and improvement of generated speech

[1138] Subject: Server

[1139] The server evaluates the quality of the generated voice to test new voices created by the generative AI model.

[1140] Voice quality assessment metrics are used in the testing.

[1141] Based on the test results, the parameters of the AI ​​model are readjusted and retrained.

[1142] NFT assignment for original audio

[1143] Subject: Server

[1144] The server generates an NFT for the voice actor's original voice.

[1145] The generated NFT is registered on the blockchain, proving ownership and authenticity of the audio data.

[1146] Audio Usage and Revenue Sharing

[1147] Subject: Terminal (animation production company)

[1148] The animation production company's terminal downloads the generated audio data and NFT from the server.

[1149] The downloaded audio data will be used in the production of anime works.

[1150] After the anime work is released, revenue data is uploaded from the terminal to the server.

[1151] Subject: Server

[1152] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[1153] An electronic payment system will be used for revenue sharing.

[1154] Audience emotion recognition and feedback collection

[1155] Subject: User (viewer)

[1156] Users can view the completed animation and enter their impressions in a feedback form.

[1157] The emotion engine recognizes and collects audience emotional data.

[1158] Subject: Server

[1159] The server analyzes the feedback from viewers and the emotional data from the emotion engine.

[1160] The analysis results will be used to further improve the AI ​​model, contributing to improving the quality of generated speech.

[1161] Examples and prompts

[1162] For example, this system would be effective if a voice actor playing a key character in an anime had to take a break due to illness. The voice data of the voice actor recorded on the device is uploaded to a server and analyzed. A generative AI model is trained based on the analysis results, and a new voice is generated. The original voice is assigned an NFT and registered on the blockchain. The animation production company downloads the generated voice and completes the work. Revenue is returned to the voice actor, and viewer emotional data and feedback are used to improve the AI ​​model.

[1163] Example prompt sentence:

[1164] "Generate the following line in the voice of the specified voice actor. Line: 'Hello, this is Character A. It's a great day today.'"

[1165] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1166] Divide the processing flow of this system's program into processing steps

[1167] Step 1: Collect and upload audio data

[1168] Step 2: Analyzing the audio data

[1169] Step 3: Training the generative AI model

[1170] Step 4: Evaluate and improve the generated speech

[1171] Step 5: Adding an NFT to the original audio

[1172] Step 6: Download and use the generated audio

[1173] Step 7: Revenue sharing

[1174] Step 8: Recognize audience emotions and collect feedback

[1175] Description of each processing step

[1176] Step 1: Collect and upload audio data

[1177] Subject: Terminal

[1178] The device uses a high-quality microphone to record voice data from voice actors, including the character's lines and various emotional expressions.

[1179] The device temporarily stores the recorded audio data and prepares it for uploading to the server.

[1180] Upload the audio file to the server using a secure communication protocol (e.g. HTTPS).

[1181] Input: Recorded audio file

[1182] Output: Audio file uploaded to the server

[1183] Step 2: Analyzing the audio data

[1184] Subject: Server

[1185] The server launches voice analysis software to analyze the received voice data.

[1186] The audio data is decomposed into features such as waveform, spectrum, pitch, volume, and intonation, and these data are generated.

[1187] The server stores the generated feature data in a database.

[1188] Input: Uploaded audio file

[1189] Output: Audio feature data

[1190] Step 3: Training the generative AI model

[1191] Subject: Server

[1192] The server uses the analyzed audio feature data to train a generative AI model.

[1193] Optimize the parameters of AI models using deep learning frameworks (e.g., TensorFlow, PyTorch).

[1194] The training process requires a lot of data and computational resources to reproduce the unique texture of the voice actor.

[1195] Input: Audio feature data

[1196] Output: A trained generative AI model

[1197] Step 4: Evaluate and improve the generated speech

[1198] Subject: Server

[1199] The server uses a generative AI model to generate new speech and evaluate its quality.

[1200] We use voice quality evaluation metrics (e.g., PESQ, STOI) and improve the AI ​​model based on the evaluation results.

[1201] The server repeats this process to improve the accuracy of the model.

[1202] Input: A trained generative AI model

[1203] Output: Evaluated generated speech data

[1204] Step 5: Adding an NFT to the original audio

[1205] Subject: Server

[1206] The server generates an NFT (non-fungible token) for the voice actor's original voice data.

[1207] NFTs are created and registered on a blockchain platform (e.g., Ethereum).

[1208] The generated NFT proves ownership and authenticity of the audio data.

[1209] Input: Original audio data

[1210] Output: NFT registered on the blockchain

[1211] Step 6: Download and use the generated audio

[1212] Subject: Terminal (animation production company)

[1213] The animation production company's terminal downloads the generated audio data and NFT from the server.

[1214] The device uses the downloaded audio data and applies it to the animation work being produced.

[1215] Input: Generated audio data, NFT

[1216] Output: Audio data and NFTs used as part of the anime production

[1217] Step 7: Revenue sharing

[1218] Subject: Terminal (animation production company)

[1219] The device records revenue data for released anime works.

[1220] Periodically upload revenue data to a server.

[1221] Subject: Server

[1222] The server receives the revenue data and distributes the revenue to the voice actors based on pre-set percentages.

[1223] Revenue sharing will be via electronic payment systems (e.g., online payment services).

[1224] Input: Revenue Data

[1225] Output: Revenue share to voice actors

[1226] Step 8: Recognize audience emotions and collect feedback

[1227] Subject: User (viewer)

[1228] Users can view the completed animation and enter their impressions in a feedback form.

[1229] The emotion engine recognizes and collects audience emotional data.

[1230] Input: Viewer sentiment data and feedback

[1231] Output: Emotion data and feedback collected on the server

[1232] Subject: Server

[1233] The server analyzes the feedback from viewers and the emotional data collected by the emotion engine.

[1234] The analysis results are used to improve the AI ​​model, leading to improved quality of the generated voice.

[1235] Input: Collected emotion data and feedback

[1236] Output: An improved generative AI model

[1237] (Application example 2)

[1238] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1239] In anime production, if a key voice actor is temporarily unavailable, it is difficult to complete the work while maintaining the quality of that character's voice. Furthermore, while viewer feedback is important for improving the quality of voice data generated by generative AI models, traditional feedback collection methods have the problem of being unable to recognize user emotions in detail. Furthermore, there is a need for a mechanism to guarantee the authenticity and ownership of generated voices and to appropriately distribute revenue.

[1240] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice analysis means, a means for training the generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion recognition means, and a means for analyzing the collected emotion data and improving the generative AI model based on the analysis results. This makes it possible to generate high-quality voice even when the main voice actor is unavailable, collect feedback based on viewer emotions, improve the generative AI model, and improve the quality of the generated voice. Furthermore, assigning NFTs to the generated voice guarantees the authenticity and ownership of the voice data and enables appropriate revenue distribution.

[1241] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[1242] "Means for training a generative AI model" means means for training a generative AI model based on analyzed audio data.

[1243] "Means for evaluating generated voice" refers to a means for evaluating the quality of voice generated by the generation AI.

[1244] "Means for assigning NFTs to audio data" refers to a means for assigning NFTs to generated audio data and registering it on the blockchain.

[1245] The "means for downloading generated voice" is a means for downloading generated voice data.

[1246] "Revenue distribution means" means a means for recording and managing revenues from generated voices and animation works, and distributing revenues to voice actors and related parties.

[1247] The "means for collecting feedback from viewers" is a means for collecting feedback from viewers and analyzing the data.

[1248] An "emotion recognition means" is a means for recognizing the viewer's emotions and collecting data on them.

[1249] "Means for analyzing collected emotional data and improving the generative AI model based on the results of the analysis" refers to means for analyzing emotional data collected from viewers and improving the generative AI model based on the results of the analysis.

[1250] The system of the present invention is realized by combining various means as follows. Specific embodiments are shown below.

[1251] 1. Collection and analysis of audio data

[1252] Subject: Terminal

[1253] Voice actors' voice data is recorded using high-quality recording equipment, and this voice includes the characters' lines and various emotional expressions.

[1254] The recorded audio data is uploaded from the device to a server, using high security and data compression technology to prevent data loss.

[1255] Subject: Server

[1256] The server receives the uploaded audio data and analyzes characteristics such as the audio waveform, spectrum, pitch, and tempo. This analysis is performed using libraries such as "librosa."

[1257] The analyzed audio feature data is stored in a database for use in training subsequent generative AI models.

[1258] 2. Training a generative AI model

[1259] Subject: Server

[1260] The server trains a generative AI model based on the speech feature data for analysis, using machine learning frameworks such as the "transformers" library.

[1261] The model is tuned to reproduce the unique texture of a particular voice actor's voice.

[1262] 3. Evaluation of generated speech

[1263] Subject: Server

[1264] The server generates new audio using the trained generative AI model and evaluates its quality, applying criteria for voice quality assessment and further fine-tuning the model if necessary.

[1265] 4. NFT assignment to original audio

[1266] Subject: Server

[1267] The server assigns an NFT to the voice actor's original voice data and registers the NFT on the blockchain, using an external blockchain API.

[1268] A registered NFT serves as proof of ownership and authenticity of the audio data.

[1269] 5. Audio Use and Revenue Sharing

[1270] Subject: Terminal (animation production company)

[1271] The animation production company downloads the generated audio data and NFTs from the server using a dedicated interface.

[1272] The downloaded generated audio is used to create an animated work, and after the work is released, revenue data is recorded and sent to a server.

[1273] Subject: Server

[1274] The server returns the revenue to the voice actor based on the received revenue data at a preset rate, using electronic payment technology.

[1275] 6. Audience Emotion Recognition and Feedback Collection

[1276] Subject: User (viewer)

[1277] Users can watch the finished animation and provide feedback on the quality of the generated voice, while the system recognizes the user's emotions in real time using the "transformers" library.

[1278] Subject: Server

[1279] The server collects and analyzes viewer feedback and emotion recognition results, which are used to improve the next generation AI model.

[1280] Specific examples

[1281] For example, there may be an anime featuring a character played by a well-known voice actor, but that voice actor is suddenly unable to appear. In this case, previously recorded voice data expressing the actor's lines and emotions is uploaded to a server, and a generative AI model is trained based on that data. Furthermore, by entering a prompt sentence as follows, character voices for specific situations can be generated.

[1282] (Example prompt): "Good morning, let's do our best today!" (in a cheerful tone)

[1283] This voice is used in the animation, and viewers can provide emotional feedback, such as "I'm excited." This feedback is analyzed in detail using emotion recognition technology and used to further improve the quality of the generative AI model.

[1284] In this way, the system of the present invention can provide high-quality generated audio that takes into account the emotions of the viewer, allowing for smooth animation production.

[1285] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1286] Step 1:

[1287] Subject: Terminal

[1288] High-quality recording equipment is used to record voice data from voice actors, including the characters' lines and various emotional expressions.

[1289] The recorded audio data is uploaded from the device to the server, using high security and data compression technology to prevent data loss.

[1290] Input: Voice actor's voice data

[1291] Output: Uploaded audio data

[1292] Step 2:

[1293] Subject: Server

[1294] The server receives the uploaded audio data.

[1295] The received audio data is analyzed and features such as audio waveform, spectrum, pitch, and tempo are extracted.

[1296] The extracted feature data is stored in a database. The "librosa" library is used to extract audio features.

[1297] Input: Uploaded audio data

[1298] Output: Analyzed audio feature data

[1299] Step 3:

[1300] Subject: Server

[1301] A generative AI model is trained based on the analyzed audio feature data using a machine learning framework such as the "transformers" library.

[1302] The trained generative AI model is tuned to reproduce the unique texture of a particular voice actor's voice.

[1303] Input: Analyzed speech feature data

[1304] Output: A trained generative AI model

[1305] Step 4:

[1306] Subject: Server

[1307] It uses trained generative AI models to generate new voices, using prompts to tailor voices to specific tones and situations.

[1308] Example prompt: "Good morning, let's do our best today!" (in a cheerful tone)

[1309] Input: Trained generative AI model, prompt

[1310] Output: Generated audio data

[1311] Step 5:

[1312] Subject: Server

[1313] The quality of the generated speech is evaluated, applying criteria for speech quality assessment and further fine-tuning the model if necessary.

[1314] Input: Generated audio data

[1315] Output: Evaluation results, improved generative AI model

[1316] Step 6:

[1317] Subject: Server

[1318] The generated audio data is assigned an NFT and registered on the blockchain using an external blockchain API.

[1319] Input: Generated audio data

[1320] Output: Audio data with NFT attached

[1321] Step 7:

[1322] Subject: Terminal (animation production company)

[1323] Download the generated audio data and NFT from the server using a dedicated interface.

[1324] Create an animated work using the downloaded generated audio.

[1325] Input: Audio data with NFT attached

[1326] Output: Anime work

[1327] Step 8:

[1328] Subject: Terminal (animation production company)

[1329] After the anime work is released, revenue data is recorded and sent to a server.

[1330] Input: Revenue Data

[1331] Output: Revenue data sent to the server

[1332] Step 9:

[1333] Subject: Server

[1334] Based on the received revenue data, revenue is returned to the voice actor based on a preset percentage, and electronic payment technology is used here.

[1335] Input: Revenue Data

[1336] Output: Revenue share to voice actors

[1337] Step 10:

[1338] Subject: User (viewer)

[1339] View the finished animation and provide feedback on the quality of the generated audio.

[1340] The emotions of the user are recognized in real time while watching. The "transformers" library is used for emotion recognition.

[1341] Input: Viewer feedback, emotion recognition data

[1342] Output: Feedback data, emotion recognition data

[1343] Step 11:

[1344] Subject: Server

[1345] Collect and analyze viewer feedback and emotion recognition results.

[1346] The analysis results will be used to improve the next generative AI model.

[1347] Input: Feedback data, emotion recognition data

[1348] Output: Analysis results, improved generative AI model

[1349] The above are the specific processing steps for carrying out the present invention. The specific operations performed in each step, the software used, and the data processing process have been shown.

[1350] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1351] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1352] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1353] [Third embodiment]

[1354] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1355] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1356] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1357] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1358] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1359] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1360] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1361] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1362] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1363] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1364] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1365] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1366] overview

[1367] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns NFTs to them to preserve the value of the originals. This aims to maintain the quality of anime works even when voice actors are unable to perform, and to stabilize voice actors' earnings.

[1368] Overall system configuration

[1369] The system of the present invention comprises the following elements:

[1370] 1. Voice analysis method: Analyze the voice actor's voice data.

[1371] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[1372] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[1373] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[1374] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[1375] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[1376] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[1377] Explanation of system processing

[1378] 1. Collection and analysis of audio data

[1379] Subject: Terminal

[1380] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[1381] The device uploads the recorded audio data to the server.

[1382] Subject: Server

[1383] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[1384] 2. Training a generative AI model

[1385] Subject: Server

[1386] The server uses the analyzed audio feature data to train a generative AI model.

[1387] This AI model is designed to reproduce the unique vocal qualities of voice actors.

[1388] 3. Evaluation of generated speech

[1389] Subject: Server

[1390] The server generates new voices using generative AI and tests them to evaluate their quality.

[1391] The model is repeatedly improved based on the evaluation results.

[1392] 4. NFT assignment to original audio

[1393] Subject: Server

[1394] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[1395] Registered NFTs prove ownership and authenticity of audio data.

[1396] 5. Audio Use and Revenue Sharing

[1397] Subject: Terminal (animation production company)

[1398] The device downloads the generated audio data and NFT from the server.

[1399] Use synthetic audio in animation production and release the finished work.

[1400] The revenue of the work is recorded and sent to the server.

[1401] Subject: Server

[1402] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[1403] 6. Quality check and feedback

[1404] Subject: User (viewer)

[1405] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[1406] Subject: Server

[1407] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[1408] Specific examples

[1409] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This system eliminates the decline in quality and instability of revenue that can occur when voice actors are replaced.

[1410] The processing flow will be explained below.

[1411] Step 1:

[1412] Audio data collection

[1413] Subject: Terminal

[1414] The device records voice actors' voices in high quality, including the characters' lines and various emotional expressions.

[1415] Step 2:

[1416] Uploading audio data

[1417] Subject: Terminal

[1418] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure secure transmission.

[1419] Step 3:

[1420] Analysis of audio data

[1421] Subject: Server

[1422] The server analyzes the received audio data. This analysis involves extracting features such as the waveform, spectrum, pitch, volume, and intonation of the audio. The analysis results are stored in a database.

[1423] Step 4:

[1424] Training generative AI models

[1425] Subject: Server

[1426] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[1427] Step 5:

[1428] Generating synthetic speech

[1429] Subject: Server

[1430] The server uses a trained generative AI model to generate new voice samples, which are then compared to the original voice.

[1431] Step 6:

[1432] Evaluation of generated speech

[1433] Subject: Server

[1434] The server evaluates the quality of the generated voice, including voice similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[1435] Step 7:

[1436] NFT granting

[1437] Subject: Server

[1438] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the voice data.

[1439] Step 8:

[1440] Download generated audio

[1441] Subject: Terminal (animation production company)

[1442] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[1443] Step 9:

[1444] Anime production and revenue records

[1445] Subject: Terminal

[1446] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[1447] Step 10:

[1448] Revenue sharing

[1449] Subject: Server

[1450] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payments.

[1451] Step 11:

[1452] Gathering feedback from viewers

[1453] Subject: User

[1454] Users can watch the completed animation and provide feedback to the platform on the quality of the generated audio and their impressions.

[1455] Step 12:

[1456] Feedback analysis and AI model improvement

[1457] Subject: Server

[1458] The server collects and analyzes feedback from viewers, and the analysis results are used to further refine the AI ​​model, thereby improving the quality of the generated voice in future iterations.

[1459] Example 1

[1460] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1461] If a voice actor is unable to provide voice over for illness or other reasons, the quality of the anime work declines and revenue becomes unstable. Furthermore, there is no way to guarantee the authenticity of the original voice over, which creates the risk of counterfeiting or unauthorized use. Therefore, a method is needed to effectively collect viewer feedback and improve generative AI models.

[1462] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1463] In this invention, the server includes a voice analysis means for analyzing voice data and extracting feature data; a generative AI model training means for training a generative AI model based on the analyzed feature data; a generated voice evaluation means for generating new voices using the generative AI model and evaluating their quality; a voice data NFT assignment means for assigning NFTs to the voice actor's original voice data and registering it on the blockchain; a revenue distribution means for returning revenue to the voice actor based on revenue data; and a feedback collection means for collecting feedback from viewers and using it to improve the generative AI model. This enables the production of high-quality anime works even when voice actors are unavailable, improving revenue stability. It also ensures the authenticity of the original voice and enables continuous improvement of the AI ​​model through feedback.

[1464] 1. "Audio data" means data that is a digital recording of the sounds and lines spoken by a voice actor.

[1465] 2. "Terminal means" means a hardware device or software tool for recording or uploading audio data.

[1466] 3. "Server" refers to a computer system that performs processes such as analyzing voice data, evaluating generated voice data, and assigning NFTs.

[1467] 4. "Audio analysis means" means technologies or algorithms that process uploaded audio data to extract characteristics such as waveform, spectrum, pitch, volume, and intonation.

[1468] 5. A "generative AI model" is a machine learning model that generates new speech based on speech data.

[1469] 6. "Generative AI model training means" means the process of optimizing and training a generative AI model using analyzed speech feature data.

[1470] 7. “Generative Speech Evaluation Measures” means tests or criteria for assessing the quality of speech generated by a generative AI model.

[1471] 8. "NFT" stands for "Non-Fungible Token" and is a token used to prove ownership and authenticity of digital assets.

[1472] 9. "Audio data NFT assignment method" is a technology that generates NFTs for audio data and registers them on the blockchain.

[1473] 10. "Revenue sharing mechanism" is a system for recording revenue from anime works and distributing that revenue fairly to voice actors.

[1474] 11. “Feedback collection methods” are technologies and methods used to collect audience ratings and feedback and use it to improve the generative AI model.

[1475] MODE FOR CARRYING OUT THE INVENTION

[1476] This invention is a system that uses generative AI technology to replicate the voice of a voice actor and assigns an NFT to preserve the value of the original. Specific embodiments for implementing this system are described below.

[1477] 1. Hardware Configuration

[1478] Terminal

[1479] Use high-quality dedicated microphones and studio equipment to record audio data, and a computer or dedicated device to capture the recorded audio data in digital format (e.g., WAV files) and upload it to a server via an internet connection.

[1480] server

[1481] It uses a high-performance server that analyzes voice data, trains generative AI models, evaluates generated voices, assigns NFTs, distributes revenue, and collects feedback. Specifically, it refers to a server with the computing resources to run libraries such as Python, TensorFlow, and PyTorch.

[1482] 2. Software Configuration

[1483] Voice analysis methods

[1484] We use LibROSA or other audio analysis libraries to extract features from the audio data, such as waveform, spectrum, pitch, volume, and intonation, and store these in a database.

[1485] Generative AI model training tools

[1486] To train the generative AI model, we use Python and machine learning libraries such as TensorFlow and PyTorch. Using the analyzed audio feature data, we optimize the model (e.g., WaveNet, Tacotron2) to reproduce the unique voice quality of the voice actor.

[1487] Generated speech evaluation means

[1488] To evaluate the quality of the generated speech, we evaluate the performance of the model using speech evaluation criteria such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[1489] Audio data NFT granting method

[1490] To assign an NFT to audio data, a smart contract creation tool (e.g., Solidity) is used and registered on the Ethereum blockchain to prove the authenticity and ownership of the audio data.

[1491] Revenue sharing method

[1492] Revenue data is collected through online payment systems (e.g., PayPal, Stripe), and revenue is distributed to voice actors based on that data. Revenue distribution is calculated using automated scripts and programs on the server.

[1493] Feedback collection methods

[1494] A dedicated feedback form and application are used to collect feedback from viewers, and the collected data is used as an evaluation index for the generative AI model, helping to improve the model.

[1495] Specific examples

[1496] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. The original voice is assigned an NFT to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This process prevents a decline in quality due to the voice actor's absence.

[1497] Prompt Sentence Examples

[1498] "As a first step, please record voice actor A's past voice data and upload it to our server. Next, we will use the server to analyze this voice data and train a generative AI model. Once the model is trained, we will evaluate the quality of the generated voice and make improvements if necessary. Finally, please attach an NFT to the original voice and download the generated voice for use."

[1499] This invention makes it possible to produce high-quality anime even when voice actors are unavailable, improving revenue stability and enabling continuous improvement of AI models based on viewer feedback.

[1500] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1501] Step 1: Record and upload audio data

[1502] Subject: Terminal

[1503] The device uses dedicated microphones and studio equipment to record high-quality voice data from voice actors. The recorded voice data (e.g., in WAV format) includes the character's lines and emotional expressions.

[1504] Once recording is complete, the device uploads the audio data to a server over the internet using a secure file transfer protocol (e.g., SFTP).

[1505] Input: Voice actor's voice data

[1506] Output: Upload audio data to the server

[1507] Step 2: Analyzing the audio data

[1508] Subject: Server

[1509] The server receives the uploaded audio data and uses an audio analysis library such as LibROSA to extract features such as waveform, spectrum, pitch, volume, and intonation.

[1510] The analyzed voice feature data is stored in a database.

[1511] Input: Uploaded audio data

[1512] Output: Analyzed audio feature data

[1513] Step 3: Training the generative AI model

[1514] Subject: Server

[1515] The server trains a generative AI model based on the analyzed voice feature data, using Python and machine learning libraries such as TensorFlow and PyTorch. The generative AI model employs a multilayer neural network-based model such as WaveNet or Tacotron2.

[1516] The model is trained to replicate the voice actor's unique vocal timbre and speaking style.

[1517] Input: Analyzed audio feature data

[1518] Output: A trained generative AI model

[1519] Step 4: Evaluate and improve the generated speech

[1520] Subject: Server

[1521] The server generates test audio using a trained generative AI model, which is then evaluated using audio metrics such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[1522] Based on the evaluation results, improvements are made repeatedly by adjusting the model's hyperparameters and retraining with additional data.

[1523] Input: Trained generative AI model, test audio data

[1524] Output: Evaluation results, improved generative AI model

[1525] Step 5: Adding an NFT to the original audio

[1526] Subject: Server

[1527] The server generates an NFT for the original audio data, using a smart contract to prove ownership and authenticity of the audio data.

[1528] The generated NFT is registered on the Ethereum blockchain and acts as a digital certificate.

[1529] Input: Original audio data

[1530] Output: NFT, registered on the blockchain

[1531] Step 6: Download and use the audio data

[1532] Subject: Terminal (animation production company)

[1533] The device downloads the generated audio data and NFT from the server, which is done securely using SFTP.

[1534] The animation production company will use the downloaded generated audio to create and release an animated work.

[1535] Input: Generated audio data, NFT

[1536] Output: Finished animation

[1537] Step 7: Revenue sharing

[1538] Subject: Terminal (animation production company)

[1539] The terminal records the revenue generated by the published anime works and transmits the revenue data to the server, which collects the data through online payment systems (e.g., PayPal, Stripe).

[1540] Input: Revenue data for completed anime works

[1541] Output: Send revenue data to the server

[1542] Subject: Server

[1543] The server distributes revenue to voice actors based on the received revenue data. Revenue distribution calculations are performed using automated scripts or programs based on pre-set percentages.

[1544] Input: Revenue data for anime works

[1545] Output: Revenue return to voice actors

[1546] Step 8: Gather viewer feedback and refine the model

[1547] Subject: User (viewer)

[1548] After watching the anime, users can provide their impressions and audio quality ratings using a dedicated feedback form or app.

[1549] Input: Viewer feedback

[1550] Output: Feedback data

[1551] Subject: Server

[1552] The server collects feedback from viewers and analyzes it as an evaluation index for machine learning. Based on the results of this analysis, the generative AI model is further improved.

[1553] Input: Feedback data

[1554] Output: An improved generative AI model

[1555] (Application example 1)

[1556] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1557] Conventional voice generation systems for voice actors have had problems maintaining the quality of voice data and stable revenue when voice actors are unable to physically participate in recording. Furthermore, there is no way to guarantee the authenticity or ownership of the generated voice data, making it difficult to effectively collect feedback from viewers and use it to improve AI models. The present invention aims to solve these problems by stably generating voice data for voice actors, ensuring revenue, and effectively utilizing feedback from viewers.

[1558] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1559] In this invention, the server includes a voice analysis means, a training means for a generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, a means for listening to the voice data through a listening application installed on a smartphone, a means for purchasing NFTs via the listening application, and a means for collecting feedback information via the listening application. This allows the quality of the generated voice and stability of revenue to be maintained even in situations where the voice actor cannot participate in recording, and enables the AI ​​model to be improved based on feedback from viewers.

[1560] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[1561] A "training method for a generative AI model" is a method for training an AI model based on analyzed voice data.

[1562] The "means for evaluating generated speech" is a means for evaluating and testing the quality of generated speech.

[1563] "Method of assigning NFTs to audio data" is a method of assigning NFTs to voice actor audio data to prove its authenticity and ownership.

[1564] The "means for downloading generated voice" is a means for making the generated voice data downloadable.

[1565] The "profit distribution means" is a means for appropriately distributing the profits from the content in which the generated audio is used.

[1566] "Means for collecting feedback from viewers" refers to means for collecting opinions and impressions from viewers.

[1567] "Means for listening to audio data through a listening application installed on a smartphone" refers to means for listening to audio data through an application installed on a smartphone.

[1568] "Means for purchasing NFTs via a viewing application" refers to means for purchasing NFTs using a viewing application.

[1569] The "means for collecting feedback information via a viewing application" refers to a means for collecting feedback information from viewers through a viewing application.

[1570] System configuration

[1571] The system of the present invention comprises the following elements:

[1572] 1. Voice analysis method: Analyze the voice actor's voice data.

[1573] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[1574] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[1575] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[1576] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[1577] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[1578] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[1579] 8. Means for listening to audio data through a listening application installed on a smartphone: A smartphone application is used as a means for a user to listen to audio data.

[1580] 9. Means for purchasing NFTs via the viewing application: A means for users to purchase NFTs via the viewing application.

[1581] 10. Means of collecting feedback information via a viewing application: Means of collecting viewer feedback information through a viewing application.

[1582] Explanation of program processing

[1583] 1. Collection and analysis of audio data

[1584] Subject: Terminal

[1585] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[1586] The device uploads the recorded audio data to the server.

[1587] Subject: Server

[1588] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and modulation, and stores the data in a database.

[1589] 2. Training a generative AI model

[1590] Subject: Server

[1591] The server uses the analyzed audio feature data to train a generative AI model, which is built using libraries such as TensorFlow and PyTorch.

[1592] 3. Evaluation of generated speech

[1593] Subject: Server

[1594] The server generates new voices using the generative AI and performs tests to evaluate their quality. The model is then repeatedly improved based on the evaluation results. A Python evaluation algorithm is used for the evaluation.

[1595] 4. NFT assignment to original audio

[1596] Subject: Server

[1597] The server generates an NFT using Ethereum for the voice actor's original voice data and registers it on the blockchain via the OpenSea API.

[1598] 5. Audio Use and Revenue Sharing

[1599] Subject: Terminal (animation production company)

[1600] The device downloads the generated audio data and NFTs from the server, uses the generated audio in animation production, and publishes the finished work. The revenue from the work is recorded and sent to the server.

[1601] Subject: Server

[1602] The server receives the revenue data and distributes the revenue to the voice actors based on a preset percentage. AWS is used as the revenue management system.

[1603] 6. Quality check and feedback

[1604] Subject: User (viewer)

[1605] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[1606] Subject: Server

[1607] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[1608] Examples of specific examples and prompts

[1609] For example, we will show a specific example using the smartphone app "Spoken NFT."

[1610] 1. App launch: The user launches the "Spoken NFT" app and searches for the generated voice of their favorite voice actor.

[1611] 2. Listen to the audio: Select the generated audio from the search results and listen to it on your smartphone.

[1612] 3. Purchase NFT: Purchase an NFT for the generated audio you like and retain ownership on the blockchain.

[1613] 4. Provide feedback: After listening, send feedback about the generated audio through the app.

[1614] Here are some example prompts to input to a generative AI model:

[1615] "Generate lines for character A with the characteristics of voice actor XX:

[1616] "From today onwards, you are a part of our team!"

[1617] Emotion: Joy

[1618] Based on this prompt, the AI ​​will reproduce the voice quality of voice actor XX and the characteristics of character A, and generate voice that also takes emotional expression into account.

[1619] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1620] Step 1: Collecting audio data

[1621] Subject: Terminal

[1622] The device records high-quality voice data from the voice actor. This voice includes the character's lines and emotional expressions. The device then uploads this voice data to the server. The input is the recorded voice data, and the output is the audio file uploaded to the server.

[1623] Step 2: Analyzing the audio data

[1624] Subject: Server

[1625] The server analyzes the received audio data by breaking it down into features such as waveform, spectrum, pitch, volume, and modulation. The analysis results are stored in a database. The input is the uploaded audio file, and the output is the analyzed audio feature data.

[1626] Step 3: Training the generative AI model

[1627] Subject: Server

[1628] The server uses the analyzed voice feature data to train a generative AI model using TensorFlow or PyTorch. During the training process, data processing and calculations are performed to update the AI ​​model. The input is the analyzed voice feature data, and the output is a trained generative AI model.

[1629] Step 4: Generate and evaluate synthetic speech

[1630] Subject: Server

[1631] The server generates new speech data using a trained generative AI model, then evaluates the quality of the generated speech and refines the AI ​​model as needed. The input is the trained generative AI model and a text prompt, and the output is the generated new speech data.

[1632] Step 5: Adding an NFT to the audio data

[1633] Subject: Server

[1634] The server generates an NFT for the generated audio data using Ethereum and registers it on the blockchain via the OpenSea API. The input is the generated audio data, and the output is the audio data with the NFT attached and a record of its ownership.

[1635] Step 6: Download the generated audio

[1636] Subject: Terminal (animation production company)

[1637] The device downloads the generated audio data and NFT from the server. The input is the audio data with the NFT attached, and the output is the downloaded audio file.

[1638] Step 7: Revenue sharing

[1639] Subject: Server

[1640] The server receives revenue data for the work and distributes the revenue to the voice actors based on a preset percentage. The input is revenue data, and the output is distributed revenue information.

[1641] Step 8: Gather feedback from your audience

[1642] Subject: User (viewer)

[1643] After listening to the generated speech through a smartphone application, the user sends feedback. The input is the feedback content, and the output is the feedback data sent to the server.

[1644] Step 9: Improve the AI ​​model with feedback

[1645] Subject: Server

[1646] The server improves the generative AI model based on the feedback information received from viewers. The input is the feedback data, and the output is an improved generative AI model.

[1647] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1648] overview

[1649] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns them NFTs to preserve the value of the originals. Furthermore, by combining it with an emotion engine that recognizes user emotions, it is possible to analyze the impact of the generated voice on the viewer and improve the AI ​​model based on that feedback. This system aims to maintain the quality of anime works even when voice actors are unable to perform, thereby stabilizing their earnings.

[1650] Overall system configuration

[1651] The system of the present invention comprises the following elements:

[1652] 1. Voice analysis method: Analyze the voice actor's voice data.

[1653] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[1654] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[1655] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[1656] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[1657] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[1658] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[1659] 8. Emotion Engine: Recognizes and analyzes viewer emotions and uses that data to improve generative AI models.

[1660] Explanation of system processing

[1661] 1. Collection and analysis of audio data

[1662] Subject: Terminal

[1663] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[1664] The device uploads the recorded audio data to the server.

[1665] Subject: Server

[1666] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[1667] 2. Training a generative AI model

[1668] Subject: Server

[1669] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[1670] 3. Evaluation of generated speech

[1671] Subject: Server

[1672] The server generates new voices using generative AI and tests them to evaluate their quality.

[1673] The AI ​​model is repeatedly improved based on the evaluation results.

[1674] 4. NFT assignment to original audio

[1675] Subject: Server

[1676] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[1677] Registered NFTs prove ownership and authenticity of audio data.

[1678] 5. Audio Use and Revenue Sharing

[1679] Subject: Terminal (animation production company)

[1680] The device downloads the generated audio data and NFT from the server through an appropriate interface.

[1681] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[1682] Subject: Server

[1683] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payment.

[1684] 6. Audience Emotion Recognition and Feedback Collection

[1685] Subject: User (viewer)

[1686] Users can watch the completed animation and provide feedback to the platform on the quality of the generated voice and their impressions, and the emotion engine will recognize the user's emotions.

[1687] Subject: Server

[1688] The server receives and analyzes feedback from viewers and emotional data collected by the emotion engine, and the analysis results are used to further improve the AI ​​model.

[1689] Specific examples

[1690] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the voice actor's past recordings on the device is uploaded to a server and analyzed there. An AI model is trained using the analyzed feature data, and the new voice generated is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work, and a portion of the revenue is returned to the voice actor. Viewers watch the work, and emotions are recognized by an emotion engine, and further feedback is provided. The server improves the AI ​​model based on the viewer's emotional data and feedback, improving the quality of the next generated voice.

[1691] In this way, the present invention enables the production of higher quality anime works by incorporating a feedback system that takes into account the emotions of viewers while maintaining the quality of the voice actors' voices. It also establishes an effective revenue model for voice actors and provides a system that allows for the stable provision of works even in unforeseen circumstances.

[1692] The processing flow will be explained below.

[1693] Step 1:

[1694] Audio data collection

[1695] Subject: Terminal

[1696] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[1697] Step 2:

[1698] Uploading audio data

[1699] Subject: Terminal

[1700] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure security.

[1701] Step 3:

[1702] Analysis of audio data

[1703] Subject: Server

[1704] The server receives the uploaded audio data and analyzes its characteristics, such as waveform, spectrum, pitch, volume, and intonation, and stores the analysis results in a database.

[1705] Step 4:

[1706] Training generative AI models

[1707] Subject: Server

[1708] The server uses the analyzed voice feature data to train a generative AI model, at which point the model acquires the voice actor's unique vocal timbre.

[1709] Step 5:

[1710] Generating synthetic speech

[1711] Subject: Server

[1712] The server uses the trained generative AI model to generate new voice samples, which are then compared and evaluated against the original voice.

[1713] Step 6:

[1714] Evaluation of generated speech

[1715] Subject: Server

[1716] The server evaluates the quality of the generated speech, including similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[1717] Step 7:

[1718] NFT granting

[1719] Subject: Server

[1720] The server generates an NFT for the voice actor's original audio data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the audio data.

[1721] Step 8:

[1722] Download generated audio

[1723] Subject: Terminal (animation production company)

[1724] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[1725] Step 9:

[1726] Anime production and revenue records

[1727] Subject: Terminal

[1728] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[1729] Step 10:

[1730] Revenue sharing

[1731] Subject: Server

[1732] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process is carried out via electronic payment.

[1733] Step 11:

[1734] Emotion Recognition and Analysis

[1735] Subject: User (viewer)

[1736] The user watches an anime work, and the emotion engine recognizes the emotion they feel. The user's emotion data is then sent to the server.

[1737] Subject: Server

[1738] The server receives and analyzes the emotion data sent by the user, and the analysis results are used to improve the generative AI model.

[1739] Step 12:

[1740] Gathering feedback

[1741] Subject: User (viewer)

[1742] Users provide feedback to the platform about the quality and impressions of the generated speech, and the feedback is sent to the server.

[1743] Step 13:

[1744] Feedback analysis and model refinement

[1745] Subject: Server

[1746] The server receives and analyzes feedback from viewers and emotion analysis data from the emotion engine. The analysis results are used to improve the generative AI model and increase the quality of the next generated voice.

[1747] Examples:

[1748] Suppose the voice actor playing main character A in a long-running anime series has to take a break due to illness. Past voice data from the voice actor is recorded on a device and uploaded to a server. The server analyzes it and trains a generative AI model to generate new voice. The generated voice is then evaluated for quality and assigned an NFT. The animation production company downloads the generated voice and produces and releases the work, with a portion of the profits going back to the voice actor. Viewers watch the work, and the emotion engine analyzes their emotions, sending feedback to the server. The server uses the analysis results to improve the AI ​​model and increase the quality of the next generated voice. This system simultaneously maintains the quality of the voice actor's voice while also improving it to match the viewer's emotions.

[1749] Example 2

[1750] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1751] In traditional anime production, voice recording by voice actors is essential, and missing voice recording directly affects the quality of the work. If a voice actor is unable to participate in recording due to illness or other reasons, the risk of production delays increases. Voice actors' earnings also tend to be unstable. Furthermore, it is difficult to generate voices that reflect the viewer's emotions, making it difficult to consistently provide high-quality voices. Our goal is to solve these issues and provide a system that can stably provide higher-quality anime works.

[1752] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1753] In this invention, the server includes a voice analysis means, a training means for the AI ​​model, a means for evaluating the generated voice, a means for assigning NFTs to the voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion engine means for recognizing and analyzing viewers' emotional data, and a testing means for evaluating the quality of the generated voice. This allows for the generation of high-quality voice even when the voice actor is unable to participate in recording, and enables the generation of voice that reflects the viewer's emotions while maintaining the quality of the work. Furthermore, by steadily returning revenue to the voice actor, economic stability can be achieved.

[1754] A "voice analysis means" is a means for recording a voice actor's voice data and breaking down and analyzing the voice into characteristics such as waveform, spectrum, pitch, volume, and intonation.

[1755] "Means for training a generative AI model" refers to means for training a generative AI model using analyzed voice feature data so that it can reproduce the unique texture and tone of a voice actor.

[1756] The "means for evaluating generated speech" is a means for evaluating whether the generated speech is of appropriate quality, and uses speech quality evaluation metrics.

[1757] "Method of assigning NFTs to audio data" is a method of generating NFTs (non-fungible tokens) for a voice actor's original audio data, registering the NFTs on the blockchain, and proving the ownership and authenticity of the audio data.

[1758] "Means for downloading generated audio" refers to the means by which animation production companies can download the generated audio data and NFTs to their own devices.

[1759] A "revenue distribution vehicle" is a means for recording revenue earned after the release of an anime work and returning it to voice actors based on a preset percentage.

[1760] The "means for collecting feedback from viewers" refers to a means for collecting and analyzing feedback and emotional data provided by viewers after watching an anime work.

[1761] The "emotion engine means" is a means for recognizing and analyzing the viewer's emotions, and the data is used to improve the generative AI model.

[1762] "Testing means for evaluating the quality of generated speech" refers to a means for conducting tests to evaluate the quality of speech generated by a trained AI model.

[1763] MODE FOR CARRYING OUT THE INVENTION

[1764] This invention is a system that uses generative AI technology to replicate a voice actor's voice and assigns an NFT to preserve the value of the original. Furthermore, this system combines an emotion engine that recognizes the user's emotions, analyzes the impact of the generated voice on the viewer, and can improve the AI ​​model based on that feedback. This system can maintain the quality of anime works and stabilize the voice actor's income even when the voice actor is unable to participate in recording.

[1765] Overall system configuration

[1766] The system of the present invention comprises the following elements:

[1767] 1. Voice analysis method: Analyze the voice actor's voice data.

[1768] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[1769] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[1770] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[1771] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[1772] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[1773] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[1774] 8. Emotion engine means: Recognize and analyze viewer emotions and use that data to improve the generative AI model.

[1775] 9. Testing method to evaluate the quality of generated speech: Evaluate the quality of speech generated by the trained AI model.

[1776] Hardware and software used

[1777] Hardware

[1778] Recording device: A high-quality microphone (e.g., Shure SM7B), a PC for recording audio (e.g., a general-purpose laptop)

[1779] Server: A high-performance server in a data center (e.g., a virtual machine in a cloud environment)

[1780] Animation production device: PC for animation production (e.g., workstation equipped with an LCD pen tablet)

[1781] software

[1782] Voice analysis software: General-purpose voice analysis tools

[1783] Generative AI Models: A General-Purpose Deep Learning Framework

[1784] NFT Blockchain Platform: A Popular Distributed Ledger Technology

[1785] Electronic payment system: Online payment service

[1786] Feedback Collection Platform: Survey Tools

[1787] Explanation of system processing

[1788] The specific processing flow of this system will be explained below.

[1789] Audio data collection and analysis

[1790] Subject: Terminal

[1791] The device records the voice actor's voice using a high-quality microphone and generates an audio file.

[1792] The collected audio files contain the characters' lines and express various emotions.

[1793] The device uploads the audio file to the server.

[1794] Subject: Server

[1795] The server uses audio analysis software to analyze the received audio files.

[1796] The server breaks down the audio into features such as waveform, spectrum, pitch, volume, and intonation, and generates analysis data.

[1797] The server stores the analyzed data in a database.

[1798] Training generative AI models

[1799] Subject: Server

[1800] The server trains a generative AI model based on the analyzed voice feature data.

[1801] The generative AI model runs on a common deep learning framework.

[1802] Once trained, the generative AI model will be able to reproduce the unique texture and tone of a voice actor.

[1803] Evaluation and improvement of generated speech

[1804] Subject: Server

[1805] The server evaluates the quality of the generated voice to test new voices created by the generative AI model.

[1806] Voice quality assessment metrics are used in the testing.

[1807] Based on the test results, the parameters of the AI ​​model are readjusted and retrained.

[1808] NFT assignment for original audio

[1809] Subject: Server

[1810] The server generates an NFT for the voice actor's original voice.

[1811] The generated NFT is registered on the blockchain, proving ownership and authenticity of the audio data.

[1812] Audio Usage and Revenue Sharing

[1813] Subject: Terminal (animation production company)

[1814] The animation production company's terminal downloads the generated audio data and NFT from the server.

[1815] The downloaded audio data will be used in the production of anime works.

[1816] After the anime work is released, revenue data is uploaded from the terminal to the server.

[1817] Subject: Server

[1818] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[1819] An electronic payment system will be used for revenue sharing.

[1820] Audience emotion recognition and feedback collection

[1821] Subject: User (viewer)

[1822] Users can view the completed animation and enter their impressions in a feedback form.

[1823] The emotion engine recognizes and collects audience emotional data.

[1824] Subject: Server

[1825] The server analyzes the feedback from viewers and the emotional data from the emotion engine.

[1826] The analysis results will be used to further improve the AI ​​model, contributing to improving the quality of generated speech.

[1827] Examples and prompts

[1828] For example, this system would be effective if a voice actor playing a key character in an anime had to take a break due to illness. The voice data of the voice actor recorded on the device is uploaded to a server and analyzed. A generative AI model is trained based on the analysis results, and a new voice is generated. The original voice is assigned an NFT and registered on the blockchain. The animation production company downloads the generated voice and completes the work. Revenue is returned to the voice actor, and viewer emotional data and feedback are used to improve the AI ​​model.

[1829] Example prompt sentence:

[1830] "Generate the following line in the voice of the specified voice actor. Line: 'Hello, this is Character A. It's a great day today.'"

[1831] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1832] Divide the processing flow of this system's program into processing steps

[1833] Step 1: Collect and upload audio data

[1834] Step 2: Analyzing the audio data

[1835] Step 3: Training the generative AI model

[1836] Step 4: Evaluate and improve the generated speech

[1837] Step 5: Adding an NFT to the original audio

[1838] Step 6: Download and use the generated audio

[1839] Step 7: Revenue sharing

[1840] Step 8: Recognize audience emotions and collect feedback

[1841] Description of each processing step

[1842] Step 1: Collect and upload audio data

[1843] Subject: Terminal

[1844] The device uses a high-quality microphone to record voice data from voice actors, including the character's lines and various emotional expressions.

[1845] The device temporarily stores the recorded audio data and prepares it for uploading to the server.

[1846] Upload the audio file to the server using a secure communication protocol (e.g. HTTPS).

[1847] Input: Recorded audio file

[1848] Output: Audio file uploaded to the server

[1849] Step 2: Analyzing the audio data

[1850] Subject: Server

[1851] The server launches voice analysis software to analyze the received voice data.

[1852] The audio data is decomposed into features such as waveform, spectrum, pitch, volume, and intonation, and these data are generated.

[1853] The server stores the generated feature data in a database.

[1854] Input: Uploaded audio file

[1855] Output: Audio feature data

[1856] Step 3: Training the generative AI model

[1857] Subject: Server

[1858] The server uses the analyzed audio feature data to train a generative AI model.

[1859] Optimize the parameters of AI models using deep learning frameworks (e.g., TensorFlow, PyTorch).

[1860] The training process requires a lot of data and computational resources to reproduce the unique texture of the voice actor.

[1861] Input: Audio feature data

[1862] Output: A trained generative AI model

[1863] Step 4: Evaluate and improve the generated speech

[1864] Subject: Server

[1865] The server uses a generative AI model to generate new speech and evaluate its quality.

[1866] We use voice quality evaluation metrics (e.g., PESQ, STOI) and improve the AI ​​model based on the evaluation results.

[1867] The server repeats this process to improve the accuracy of the model.

[1868] Input: A trained generative AI model

[1869] Output: Evaluated generated speech data

[1870] Step 5: Adding an NFT to the original audio

[1871] Subject: Server

[1872] The server generates an NFT (non-fungible token) for the voice actor's original voice data.

[1873] NFTs are created and registered on a blockchain platform (e.g., Ethereum).

[1874] The generated NFT proves ownership and authenticity of the audio data.

[1875] Input: Original audio data

[1876] Output: NFT registered on the blockchain

[1877] Step 6: Download and use the generated audio

[1878] Subject: Terminal (animation production company)

[1879] The animation production company's terminal downloads the generated audio data and NFT from the server.

[1880] The device uses the downloaded audio data and applies it to the animation work being produced.

[1881] Input: Generated audio data, NFT

[1882] Output: Audio data and NFTs used as part of the anime production

[1883] Step 7: Revenue sharing

[1884] Subject: Terminal (animation production company)

[1885] The device records revenue data for released anime works.

[1886] Periodically upload revenue data to a server.

[1887] Subject: Server

[1888] The server receives the revenue data and distributes the revenue to the voice actors based on pre-set percentages.

[1889] Revenue sharing will be via electronic payment systems (e.g., online payment services).

[1890] Input: Revenue Data

[1891] Output: Revenue share to voice actors

[1892] Step 8: Recognize audience emotions and collect feedback

[1893] Subject: User (viewer)

[1894] Users can view the completed animation and enter their impressions in a feedback form.

[1895] The emotion engine recognizes and collects audience emotional data.

[1896] Input: Viewer sentiment data and feedback

[1897] Output: Emotion data and feedback collected on the server

[1898] Subject: Server

[1899] The server analyzes the feedback from viewers and the emotional data collected by the emotion engine.

[1900] The analysis results are used to improve the AI ​​model, leading to improved quality of the generated voice.

[1901] Input: Collected emotion data and feedback

[1902] Output: An improved generative AI model

[1903] (Application example 2)

[1904] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1905] In anime production, if a key voice actor is temporarily unavailable, it is difficult to complete the work while maintaining the quality of that character's voice. Furthermore, while viewer feedback is important for improving the quality of voice data generated by generative AI models, traditional feedback collection methods have the problem of being unable to recognize user emotions in detail. Furthermore, there is a need for a mechanism to guarantee the authenticity and ownership of generated voices and to appropriately distribute revenue.

[1906] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice analysis means, a means for training the generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion recognition means, and a means for analyzing the collected emotion data and improving the generative AI model based on the analysis results. This makes it possible to generate high-quality voice even when the main voice actor is unavailable, collect feedback based on viewer emotions, improve the generative AI model, and improve the quality of the generated voice. Furthermore, assigning NFTs to the generated voice guarantees the authenticity and ownership of the voice data and enables appropriate revenue distribution.

[1907] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[1908] "Means for training a generative AI model" means means for training a generative AI model based on analyzed audio data.

[1909] "Means for evaluating generated voice" refers to a means for evaluating the quality of voice generated by the generation AI.

[1910] "Means for assigning NFTs to audio data" refers to a means for assigning NFTs to generated audio data and registering it on the blockchain.

[1911] The "means for downloading generated voice" is a means for downloading generated voice data.

[1912] "Revenue distribution means" means a means for recording and managing revenues from generated voices and animation works, and distributing revenues to voice actors and related parties.

[1913] The "means for collecting feedback from viewers" is a means for collecting feedback from viewers and analyzing the data.

[1914] An "emotion recognition means" is a means for recognizing the viewer's emotions and collecting data on them.

[1915] "Means for analyzing collected emotional data and improving the generative AI model based on the results of the analysis" refers to means for analyzing emotional data collected from viewers and improving the generative AI model based on the results of the analysis.

[1916] The system of the present invention is realized by combining various means as follows. Specific embodiments are shown below.

[1917] 1. Collection and analysis of audio data

[1918] Subject: Terminal

[1919] Voice actors' voice data is recorded using high-quality recording equipment, and this voice includes the characters' lines and various emotional expressions.

[1920] The recorded audio data is uploaded from the device to a server, using high security and data compression technology to prevent data loss.

[1921] Subject: Server

[1922] The server receives the uploaded audio data and analyzes characteristics such as the audio waveform, spectrum, pitch, and tempo. This analysis is performed using libraries such as "librosa."

[1923] The analyzed audio feature data is stored in a database for use in training subsequent generative AI models.

[1924] 2. Training a generative AI model

[1925] Subject: Server

[1926] The server trains a generative AI model based on the speech feature data for analysis, using machine learning frameworks such as the "transformers" library.

[1927] The model is tuned to reproduce the unique texture of a particular voice actor's voice.

[1928] 3. Evaluation of generated speech

[1929] Subject: Server

[1930] The server generates new audio using the trained generative AI model and evaluates its quality, applying criteria for voice quality assessment and further fine-tuning the model if necessary.

[1931] 4. NFT assignment to original audio

[1932] Subject: Server

[1933] The server assigns an NFT to the voice actor's original voice data and registers the NFT on the blockchain, using an external blockchain API.

[1934] A registered NFT serves as proof of ownership and authenticity of the audio data.

[1935] 5. Audio Use and Revenue Sharing

[1936] Subject: Terminal (animation production company)

[1937] The animation production company downloads the generated audio data and NFTs from the server using a dedicated interface.

[1938] The downloaded generated audio is used to create an animated work, and after the work is released, revenue data is recorded and sent to a server.

[1939] Subject: Server

[1940] The server returns the revenue to the voice actor based on the received revenue data at a preset rate, using electronic payment technology.

[1941] 6. Audience Emotion Recognition and Feedback Collection

[1942] Subject: User (viewer)

[1943] Users can watch the finished animation and provide feedback on the quality of the generated voice, while the system recognizes the user's emotions in real time using the "transformers" library.

[1944] Subject: Server

[1945] The server collects and analyzes viewer feedback and emotion recognition results, which are used to improve the next generation AI model.

[1946] Specific examples

[1947] For example, there may be an anime featuring a character played by a well-known voice actor, but that voice actor is suddenly unable to appear. In this case, previously recorded voice data expressing the actor's lines and emotions is uploaded to a server, and a generative AI model is trained based on that data. Furthermore, by entering a prompt sentence as follows, character voices for specific situations can be generated.

[1948] (Example prompt): "Good morning, let's do our best today!" (in a cheerful tone)

[1949] This voice is used in the animation, and viewers can provide emotional feedback, such as "I'm excited." This feedback is analyzed in detail using emotion recognition technology and used to further improve the quality of the generative AI model.

[1950] In this way, the system of the present invention can provide high-quality generated audio that takes into account the emotions of the viewer, allowing for smooth animation production.

[1951] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1952] Step 1:

[1953] Subject: Terminal

[1954] High-quality recording equipment is used to record voice data from voice actors, including the characters' lines and various emotional expressions.

[1955] The recorded audio data is uploaded from the device to the server, using high security and data compression technology to prevent data loss.

[1956] Input: Voice actor's voice data

[1957] Output: Uploaded audio data

[1958] Step 2:

[1959] Subject: Server

[1960] The server receives the uploaded audio data.

[1961] The received audio data is analyzed and features such as audio waveform, spectrum, pitch, and tempo are extracted.

[1962] The extracted feature data is stored in a database. The "librosa" library is used to extract audio features.

[1963] Input: Uploaded audio data

[1964] Output: Analyzed audio feature data

[1965] Step 3:

[1966] Subject: Server

[1967] A generative AI model is trained based on the analyzed audio feature data using a machine learning framework such as the "transformers" library.

[1968] The trained generative AI model is tuned to reproduce the unique texture of a particular voice actor's voice.

[1969] Input: Analyzed speech feature data

[1970] Output: A trained generative AI model

[1971] Step 4:

[1972] Subject: Server

[1973] It uses trained generative AI models to generate new voices, using prompts to tailor voices to specific tones and situations.

[1974] Example prompt: "Good morning, let's do our best today!" (in a cheerful tone)

[1975] Input: Trained generative AI model, prompt

[1976] Output: Generated audio data

[1977] Step 5:

[1978] Subject: Server

[1979] The quality of the generated speech is evaluated, applying criteria for speech quality assessment and further fine-tuning the model if necessary.

[1980] Input: Generated audio data

[1981] Output: Evaluation results, improved generative AI model

[1982] Step 6:

[1983] Subject: Server

[1984] The generated audio data is assigned an NFT and registered on the blockchain using an external blockchain API.

[1985] Input: Generated audio data

[1986] Output: Audio data with NFT attached

[1987] Step 7:

[1988] Subject: Terminal (animation production company)

[1989] Download the generated audio data and NFT from the server using a dedicated interface.

[1990] Create an animated work using the downloaded generated audio.

[1991] Input: Audio data with NFT attached

[1992] Output: Anime work

[1993] Step 8:

[1994] Subject: Terminal (animation production company)

[1995] After the anime work is released, revenue data is recorded and sent to a server.

[1996] Input: Revenue Data

[1997] Output: Revenue data sent to the server

[1998] Step 9:

[1999] Subject: Server

[2000] Based on the received revenue data, revenue is returned to the voice actor based on a preset percentage, and electronic payment technology is used here.

[2001] Input: Revenue Data

[2002] Output: Revenue share to voice actors

[2003] Step 10:

[2004] Subject: User (viewer)

[2005] View the finished animation and provide feedback on the quality of the generated audio.

[2006] The emotions of the user are recognized in real time while watching. The "transformers" library is used for emotion recognition.

[2007] Input: Viewer feedback, emotion recognition data

[2008] Output: Feedback data, emotion recognition data

[2009] Step 11:

[2010] Subject: Server

[2011] Collect and analyze viewer feedback and emotion recognition results.

[2012] The analysis results will be used to improve the next generative AI model.

[2013] Input: Feedback data, emotion recognition data

[2014] Output: Analysis results, improved generative AI model

[2015] The above are the specific processing steps for carrying out the present invention. The specific operations performed in each step, the software used, and the data processing process have been shown.

[2016] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[2017] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2018] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[2019] [Fourth embodiment]

[2020] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[2021] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2023] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[2024] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[2025] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[2026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[2027] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[2028] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[2029] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2031] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[2032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2033] overview

[2034] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns NFTs to them to preserve the value of the originals. This aims to maintain the quality of anime works even when voice actors are unable to perform, and to stabilize voice actors' earnings.

[2035] Overall system configuration

[2036] The system of the present invention comprises the following elements:

[2037] 1. Voice analysis method: Analyze the voice actor's voice data.

[2038] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[2039] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[2040] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[2041] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[2042] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[2043] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[2044] Explanation of system processing

[2045] 1. Collection and analysis of audio data

[2046] Subject: Terminal

[2047] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[2048] The device uploads the recorded audio data to the server.

[2049] Subject: Server

[2050] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[2051] 2. Training a generative AI model

[2052] Subject: Server

[2053] The server uses the analyzed audio feature data to train a generative AI model.

[2054] This AI model is designed to reproduce the unique vocal qualities of voice actors.

[2055] 3. Evaluation of generated speech

[2056] Subject: Server

[2057] The server generates new voices using generative AI and tests them to evaluate their quality.

[2058] The model is repeatedly improved based on the evaluation results.

[2059] 4. NFT assignment to original audio

[2060] Subject: Server

[2061] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[2062] Registered NFTs prove ownership and authenticity of audio data.

[2063] 5. Audio Use and Revenue Sharing

[2064] Subject: Terminal (animation production company)

[2065] The device downloads the generated audio data and NFT from the server.

[2066] Use synthetic audio in animation production and release the finished work.

[2067] The revenue of the work is recorded and sent to the server.

[2068] Subject: Server

[2069] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[2070] 6. Quality check and feedback

[2071] Subject: User (viewer)

[2072] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[2073] Subject: Server

[2074] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[2075] Specific examples

[2076] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This system eliminates the decline in quality and instability of revenue that can occur when voice actors are replaced.

[2077] The processing flow will be explained below.

[2078] Step 1:

[2079] Audio data collection

[2080] Subject: Terminal

[2081] The device records voice actors' voices in high quality, including the characters' lines and various emotional expressions.

[2082] Step 2:

[2083] Uploading audio data

[2084] Subject: Terminal

[2085] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure secure transmission.

[2086] Step 3:

[2087] Analysis of audio data

[2088] Subject: Server

[2089] The server analyzes the received audio data. This analysis involves extracting features such as the waveform, spectrum, pitch, volume, and intonation of the audio. The analysis results are stored in a database.

[2090] Step 4:

[2091] Training generative AI models

[2092] Subject: Server

[2093] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[2094] Step 5:

[2095] Generating synthetic speech

[2096] Subject: Server

[2097] The server uses a trained generative AI model to generate new voice samples, which are then compared to the original voice.

[2098] Step 6:

[2099] Evaluation of generated speech

[2100] Subject: Server

[2101] The server evaluates the quality of the generated voice, including voice similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[2102] Step 7:

[2103] NFT granting

[2104] Subject: Server

[2105] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the voice data.

[2106] Step 8:

[2107] Download generated audio

[2108] Subject: Terminal (animation production company)

[2109] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[2110] Step 9:

[2111] Anime production and revenue records

[2112] Subject: Terminal

[2113] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[2114] Step 10:

[2115] Revenue sharing

[2116] Subject: Server

[2117] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payments.

[2118] Step 11:

[2119] Gathering feedback from viewers

[2120] Subject: User

[2121] Users can watch the completed animation and provide feedback to the platform on the quality of the generated audio and their impressions.

[2122] Step 12:

[2123] Feedback analysis and AI model improvement

[2124] Subject: Server

[2125] The server collects and analyzes feedback from viewers, and the analysis results are used to further refine the AI ​​model, thereby improving the quality of the generated voice in future iterations.

[2126] Example 1

[2127] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2128] If a voice actor is unable to provide voice over for illness or other reasons, the quality of the anime work declines and revenue becomes unstable. Furthermore, there is no way to guarantee the authenticity of the original voice over, which creates the risk of counterfeiting or unauthorized use. Therefore, a method is needed to effectively collect viewer feedback and improve generative AI models.

[2129] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2130] In this invention, the server includes a voice analysis means for analyzing voice data and extracting feature data; a generative AI model training means for training a generative AI model based on the analyzed feature data; a generated voice evaluation means for generating new voices using the generative AI model and evaluating their quality; a voice data NFT assignment means for assigning NFTs to the voice actor's original voice data and registering it on the blockchain; a revenue distribution means for returning revenue to the voice actor based on revenue data; and a feedback collection means for collecting feedback from viewers and using it to improve the generative AI model. This enables the production of high-quality anime works even when voice actors are unavailable, improving revenue stability. It also ensures the authenticity of the original voice and enables continuous improvement of the AI ​​model through feedback.

[2131] 1. "Audio data" means data that is a digital recording of the sounds and lines spoken by a voice actor.

[2132] 2. "Terminal means" means a hardware device or software tool for recording or uploading audio data.

[2133] 3. "Server" refers to a computer system that performs processes such as analyzing voice data, evaluating generated voice data, and assigning NFTs.

[2134] 4. "Audio analysis means" means technologies or algorithms that process uploaded audio data to extract characteristics such as waveform, spectrum, pitch, volume, and intonation.

[2135] 5. A "generative AI model" is a machine learning model that generates new speech based on speech data.

[2136] 6. "Generative AI model training means" means the process of optimizing and training a generative AI model using analyzed speech feature data.

[2137] 7. “Generative Speech Evaluation Measures” means tests or criteria for assessing the quality of speech generated by a generative AI model.

[2138] 8. "NFT" stands for "Non-Fungible Token" and is a token used to prove ownership and authenticity of digital assets.

[2139] 9. "Audio data NFT assignment method" is a technology that generates NFTs for audio data and registers them on the blockchain.

[2140] 10. "Revenue sharing mechanism" is a system for recording revenue from anime works and distributing that revenue fairly to voice actors.

[2141] 11. “Feedback collection methods” are technologies and methods used to collect audience ratings and feedback and use it to improve the generative AI model.

[2142] MODE FOR CARRYING OUT THE INVENTION

[2143] This invention is a system that uses generative AI technology to replicate the voice of a voice actor and assigns an NFT to preserve the value of the original. Specific embodiments for implementing this system are described below.

[2144] 1. Hardware Configuration

[2145] Terminal

[2146] Use high-quality dedicated microphones and studio equipment to record audio data, and a computer or dedicated device to capture the recorded audio data in digital format (e.g., WAV files) and upload it to a server via an internet connection.

[2147] server

[2148] It uses a high-performance server that analyzes voice data, trains generative AI models, evaluates generated voices, assigns NFTs, distributes revenue, and collects feedback. Specifically, it refers to a server with the computing resources to run libraries such as Python, TensorFlow, and PyTorch.

[2149] 2. Software Configuration

[2150] Voice analysis methods

[2151] We use LibROSA or other audio analysis libraries to extract features from the audio data, such as waveform, spectrum, pitch, volume, and intonation, and store these in a database.

[2152] Generative AI model training tools

[2153] To train the generative AI model, we use Python and machine learning libraries such as TensorFlow and PyTorch. Using the analyzed audio feature data, we optimize the model (e.g., WaveNet, Tacotron2) to reproduce the unique voice quality of the voice actor.

[2154] Generated speech evaluation means

[2155] To evaluate the quality of the generated speech, we evaluate the performance of the model using speech evaluation criteria such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[2156] Audio data NFT granting method

[2157] To assign an NFT to audio data, a smart contract creation tool (e.g., Solidity) is used and registered on the Ethereum blockchain to prove the authenticity and ownership of the audio data.

[2158] Revenue sharing method

[2159] Revenue data is collected through online payment systems (e.g., PayPal, Stripe), and revenue is distributed to voice actors based on that data. Revenue distribution is calculated using automated scripts and programs on the server.

[2160] Feedback collection methods

[2161] A dedicated feedback form and application are used to collect feedback from viewers, and the collected data is used as an evaluation index for the generative AI model, helping to improve the model.

[2162] Specific examples

[2163] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the previous voice actor recorded on a device is uploaded to a server and analyzed there. The analyzed feature data is used to train an AI model, and the generated new voice is tested. The original voice is assigned an NFT to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work and gives a portion of the revenue back to the voice actor. Viewers watch the work and provide feedback, which is used to further improve the AI ​​model. This process prevents a decline in quality due to the voice actor's absence.

[2164] Prompt Sentence Examples

[2165] "As a first step, please record voice actor A's past voice data and upload it to our server. Next, we will use the server to analyze this voice data and train a generative AI model. Once the model is trained, we will evaluate the quality of the generated voice and make improvements if necessary. Finally, please attach an NFT to the original voice and download the generated voice for use."

[2166] This invention makes it possible to produce high-quality anime even when voice actors are unavailable, improving revenue stability and enabling continuous improvement of AI models based on viewer feedback.

[2167] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2168] Step 1: Record and upload audio data

[2169] Subject: Terminal

[2170] The device uses dedicated microphones and studio equipment to record high-quality voice data from voice actors. The recorded voice data (e.g., in WAV format) includes the character's lines and emotional expressions.

[2171] Once recording is complete, the device uploads the audio data to a server over the internet using a secure file transfer protocol (e.g., SFTP).

[2172] Input: Voice actor's voice data

[2173] Output: Upload audio data to the server

[2174] Step 2: Analyzing the audio data

[2175] Subject: Server

[2176] The server receives the uploaded audio data and uses an audio analysis library such as LibROSA to extract features such as waveform, spectrum, pitch, volume, and intonation.

[2177] The analyzed voice feature data is stored in a database.

[2178] Input: Uploaded audio data

[2179] Output: Analyzed audio feature data

[2180] Step 3: Training the generative AI model

[2181] Subject: Server

[2182] The server trains a generative AI model based on the analyzed voice feature data, using Python and machine learning libraries such as TensorFlow and PyTorch. The generative AI model employs a multilayer neural network-based model such as WaveNet or Tacotron2.

[2183] The model is trained to replicate the voice actor's unique vocal timbre and speaking style.

[2184] Input: Analyzed audio feature data

[2185] Output: A trained generative AI model

[2186] Step 4: Evaluate and improve the generated speech

[2187] Subject: Server

[2188] The server generates test audio using a trained generative AI model, which is then evaluated using audio metrics such as MOS (Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality).

[2189] Based on the evaluation results, improvements are made repeatedly by adjusting the model's hyperparameters and retraining with additional data.

[2190] Input: Trained generative AI model, test audio data

[2191] Output: Evaluation results, improved generative AI model

[2192] Step 5: Adding an NFT to the original audio

[2193] Subject: Server

[2194] The server generates an NFT for the original audio data, using a smart contract to prove ownership and authenticity of the audio data.

[2195] The generated NFT is registered on the Ethereum blockchain and acts as a digital certificate.

[2196] Input: Original audio data

[2197] Output: NFT, registered on the blockchain

[2198] Step 6: Download and use the audio data

[2199] Subject: Terminal (animation production company)

[2200] The device downloads the generated audio data and NFT from the server, which is done securely using SFTP.

[2201] The animation production company will use the downloaded generated audio to create and release an animated work.

[2202] Input: Generated audio data, NFT

[2203] Output: Finished animation

[2204] Step 7: Revenue sharing

[2205] Subject: Terminal (animation production company)

[2206] The terminal records the revenue generated by the published anime works and transmits the revenue data to the server, which collects the data through online payment systems (e.g., PayPal, Stripe).

[2207] Input: Revenue data for completed anime works

[2208] Output: Send revenue data to the server

[2209] Subject: Server

[2210] The server distributes revenue to voice actors based on the received revenue data. Revenue distribution calculations are performed using automated scripts or programs based on pre-set percentages.

[2211] Input: Revenue data for anime works

[2212] Output: Revenue return to voice actors

[2213] Step 8: Gather viewer feedback and refine the model

[2214] Subject: User (viewer)

[2215] After watching the anime, users can provide their impressions and audio quality ratings using a dedicated feedback form or app.

[2216] Input: Viewer feedback

[2217] Output: Feedback data

[2218] Subject: Server

[2219] The server collects feedback from viewers and analyzes it as an evaluation index for machine learning. Based on the results of this analysis, the generative AI model is further improved.

[2220] Input: Feedback data

[2221] Output: An improved generative AI model

[2222] (Application example 1)

[2223] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2224] Conventional voice generation systems for voice actors have had problems maintaining the quality of voice data and stable revenue when voice actors are unable to physically participate in recording. Furthermore, there is no way to guarantee the authenticity or ownership of the generated voice data, making it difficult to effectively collect feedback from viewers and use it to improve AI models. The present invention aims to solve these problems by stably generating voice data for voice actors, ensuring revenue, and effectively utilizing feedback from viewers.

[2225] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2226] In this invention, the server includes a voice analysis means, a training means for a generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, a means for listening to the voice data through a listening application installed on a smartphone, a means for purchasing NFTs via the listening application, and a means for collecting feedback information via the listening application. This allows the quality of the generated voice and stability of revenue to be maintained even in situations where the voice actor cannot participate in recording, and enables the AI ​​model to be improved based on feedback from viewers.

[2227] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[2228] A "training method for a generative AI model" is a method for training an AI model based on analyzed voice data.

[2229] The "means for evaluating generated speech" is a means for evaluating and testing the quality of generated speech.

[2230] "Method of assigning NFTs to audio data" is a method of assigning NFTs to voice actor audio data to prove its authenticity and ownership.

[2231] The "means for downloading generated voice" is a means for making the generated voice data downloadable.

[2232] The "profit distribution means" is a means for appropriately distributing the profits from the content in which the generated audio is used.

[2233] "Means for collecting feedback from viewers" refers to means for collecting opinions and impressions from viewers.

[2234] "Means for listening to audio data through a listening application installed on a smartphone" refers to means for listening to audio data through an application installed on a smartphone.

[2235] "Means for purchasing NFTs via a viewing application" refers to means for purchasing NFTs using a viewing application.

[2236] The "means for collecting feedback information via a viewing application" refers to a means for collecting feedback information from viewers through a viewing application.

[2237] System configuration

[2238] The system of the present invention comprises the following elements:

[2239] 1. Voice analysis method: Analyze the voice actor's voice data.

[2240] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[2241] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[2242] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[2243] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[2244] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[2245] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[2246] 8. Means for listening to audio data through a listening application installed on a smartphone: A smartphone application is used as a means for a user to listen to audio data.

[2247] 9. Means for purchasing NFTs via the viewing application: A means for users to purchase NFTs via the viewing application.

[2248] 10. Means of collecting feedback information via a viewing application: Means of collecting viewer feedback information through a viewing application.

[2249] Explanation of program processing

[2250] 1. Collection and analysis of audio data

[2251] Subject: Terminal

[2252] The device records high-quality voice data from voice actors, including the characters' lines and emotional expressions.

[2253] The device uploads the recorded audio data to the server.

[2254] Subject: Server

[2255] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and modulation, and stores the data in a database.

[2256] 2. Training a generative AI model

[2257] Subject: Server

[2258] The server uses the analyzed audio feature data to train a generative AI model, which is built using libraries such as TensorFlow and PyTorch.

[2259] 3. Evaluation of generated speech

[2260] Subject: Server

[2261] The server generates new voices using the generative AI and performs tests to evaluate their quality. The model is then repeatedly improved based on the evaluation results. A Python evaluation algorithm is used for the evaluation.

[2262] 4. NFT assignment to original audio

[2263] Subject: Server

[2264] The server generates an NFT using Ethereum for the voice actor's original voice data and registers it on the blockchain via the OpenSea API.

[2265] 5. Audio Use and Revenue Sharing

[2266] Subject: Terminal (animation production company)

[2267] The device downloads the generated audio data and NFTs from the server, uses the generated audio in animation production, and publishes the finished work. The revenue from the work is recorded and sent to the server.

[2268] Subject: Server

[2269] The server receives the revenue data and distributes the revenue to the voice actors based on a preset percentage. AWS is used as the revenue management system.

[2270] 6. Quality check and feedback

[2271] Subject: User (viewer)

[2272] Users watch the animation and provide feedback on the quality of the generated audio and their impressions.

[2273] Subject: Server

[2274] The server receives feedback from viewers and uses the analysis to improve the generative AI model.

[2275] Examples of specific examples and prompts

[2276] For example, we will show a specific example using the smartphone app "Spoken NFT."

[2277] 1. App launch: The user launches the "Spoken NFT" app and searches for the generated voice of their favorite voice actor.

[2278] 2. Listen to the audio: Select the generated audio from the search results and listen to it on your smartphone.

[2279] 3. Purchase NFT: Purchase an NFT for the generated audio you like and retain ownership on the blockchain.

[2280] 4. Provide feedback: After listening, send feedback about the generated audio through the app.

[2281] Here are some example prompts to input to a generative AI model:

[2282] "Generate lines for character A with the characteristics of voice actor XX:

[2283] "From today onwards, you are a part of our team!"

[2284] Emotion: Joy

[2285] Based on this prompt, the AI ​​will reproduce the voice quality of voice actor XX and the characteristics of character A, and generate voice that also takes emotional expression into account.

[2286] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2287] Step 1: Collecting audio data

[2288] Subject: Terminal

[2289] The device records high-quality voice data from the voice actor. This voice includes the character's lines and emotional expressions. The device then uploads this voice data to the server. The input is the recorded voice data, and the output is the audio file uploaded to the server.

[2290] Step 2: Analyzing the audio data

[2291] Subject: Server

[2292] The server analyzes the received audio data by breaking it down into features such as waveform, spectrum, pitch, volume, and modulation. The analysis results are stored in a database. The input is the uploaded audio file, and the output is the analyzed audio feature data.

[2293] Step 3: Training the generative AI model

[2294] Subject: Server

[2295] The server uses the analyzed voice feature data to train a generative AI model using TensorFlow or PyTorch. During the training process, data processing and calculations are performed to update the AI ​​model. The input is the analyzed voice feature data, and the output is a trained generative AI model.

[2296] Step 4: Generate and evaluate synthetic speech

[2297] Subject: Server

[2298] The server generates new speech data using a trained generative AI model, then evaluates the quality of the generated speech and refines the AI ​​model as needed. The input is the trained generative AI model and a text prompt, and the output is the generated new speech data.

[2299] Step 5: Adding an NFT to the audio data

[2300] Subject: Server

[2301] The server generates an NFT for the generated audio data using Ethereum and registers it on the blockchain via the OpenSea API. The input is the generated audio data, and the output is the audio data with the NFT attached and a record of its ownership.

[2302] Step 6: Download the generated audio

[2303] Subject: Terminal (animation production company)

[2304] The device downloads the generated audio data and NFT from the server. The input is the audio data with the NFT attached, and the output is the downloaded audio file.

[2305] Step 7: Revenue sharing

[2306] Subject: Server

[2307] The server receives revenue data for the work and distributes the revenue to the voice actors based on a preset percentage. The input is revenue data, and the output is distributed revenue information.

[2308] Step 8: Gather feedback from your audience

[2309] Subject: User (viewer)

[2310] After listening to the generated speech through a smartphone application, the user sends feedback. The input is the feedback content, and the output is the feedback data sent to the server.

[2311] Step 9: Improve the AI ​​model with feedback

[2312] Subject: Server

[2313] The server improves the generative AI model based on the feedback information received from viewers. The input is the feedback data, and the output is an improved generative AI model.

[2314] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2315] overview

[2316] This invention is a system that uses generative AI technology to replicate voice actors' voices and assigns them NFTs to preserve the value of the originals. Furthermore, by combining it with an emotion engine that recognizes user emotions, it is possible to analyze the impact of the generated voice on the viewer and improve the AI ​​model based on that feedback. This system aims to maintain the quality of anime works even when voice actors are unable to perform, thereby stabilizing their earnings.

[2317] Overall system configuration

[2318] The system of the present invention comprises the following elements:

[2319] 1. Voice analysis method: Analyze the voice actor's voice data.

[2320] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[2321] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[2322] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[2323] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[2324] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[2325] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[2326] 8. Emotion Engine: Recognizes and analyzes viewer emotions and uses that data to improve generative AI models.

[2327] Explanation of system processing

[2328] 1. Collection and analysis of audio data

[2329] Subject: Terminal

[2330] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[2331] The device uploads the recorded audio data to the server.

[2332] Subject: Server

[2333] The server receives the audio data, analyzes its characteristics such as waveform, spectrum, pitch, volume, and intonation, and stores them in a database.

[2334] 2. Training a generative AI model

[2335] Subject: Server

[2336] The server uses the analyzed audio feature data to train a generative AI model, which is tuned to reproduce the unique texture of the voice actor's voice.

[2337] 3. Evaluation of generated speech

[2338] Subject: Server

[2339] The server generates new voices using generative AI and tests them to evaluate their quality.

[2340] The AI ​​model is repeatedly improved based on the evaluation results.

[2341] 4. NFT assignment to original audio

[2342] Subject: Server

[2343] The server generates an NFT for the voice actor's original voice data and registers the NFT on the blockchain.

[2344] Registered NFTs prove ownership and authenticity of audio data.

[2345] 5. Audio Use and Revenue Sharing

[2346] Subject: Terminal (animation production company)

[2347] The device downloads the generated audio data and NFT from the server through an appropriate interface.

[2348] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[2349] Subject: Server

[2350] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process uses electronic payment.

[2351] 6. Audience Emotion Recognition and Feedback Collection

[2352] Subject: User (viewer)

[2353] Users can watch the completed animation and provide feedback to the platform on the quality of the generated voice and their impressions, and the emotion engine will recognize the user's emotions.

[2354] Subject: Server

[2355] The server receives and analyzes feedback from viewers and emotional data collected by the emotion engine, and the analysis results are used to further improve the AI ​​model.

[2356] Specific examples

[2357] For example, this system would be effective if the voice actor playing long-running character A in an anime had to take a break due to illness. The voice data of the voice actor's past recordings on the device is uploaded to a server and analyzed there. An AI model is trained using the analyzed feature data, and the new voice generated is tested. An NFT is assigned to the original voice to distinguish it from the generated voice. The animation production company uses the generated voice to complete the work, and a portion of the revenue is returned to the voice actor. Viewers watch the work, and emotions are recognized by an emotion engine, and further feedback is provided. The server improves the AI ​​model based on the viewer's emotional data and feedback, improving the quality of the next generated voice.

[2358] In this way, the present invention enables the production of higher quality anime works by incorporating a feedback system that takes into account the emotions of viewers while maintaining the quality of the voice actors' voices. It also establishes an effective revenue model for voice actors and provides a system that allows for the stable provision of works even in unforeseen circumstances.

[2359] The processing flow will be explained below.

[2360] Step 1:

[2361] Audio data collection

[2362] Subject: Terminal

[2363] The device records high-quality voice data from voice actors, including the characters' lines and various emotional expressions.

[2364] Step 2:

[2365] Uploading audio data

[2366] Subject: Terminal

[2367] The device uploads the recorded audio data to a server, where it is properly encrypted to ensure security.

[2368] Step 3:

[2369] Analysis of audio data

[2370] Subject: Server

[2371] The server receives the uploaded audio data and analyzes its characteristics, such as waveform, spectrum, pitch, volume, and intonation, and stores the analysis results in a database.

[2372] Step 4:

[2373] Training generative AI models

[2374] Subject: Server

[2375] The server uses the analyzed voice feature data to train a generative AI model, at which point the model acquires the voice actor's unique vocal timbre.

[2376] Step 5:

[2377] Generating synthetic speech

[2378] Subject: Server

[2379] The server uses the trained generative AI model to generate new voice samples, which are then compared and evaluated against the original voice.

[2380] Step 6:

[2381] Evaluation of generated speech

[2382] Subject: Server

[2383] The server evaluates the quality of the generated speech, including similarity, naturalness, and accuracy of emotional expression. The evaluation results are used to improve the AI ​​model.

[2384] Step 7:

[2385] NFT granting

[2386] Subject: Server

[2387] The server generates an NFT for the voice actor's original audio data and registers the NFT on the blockchain, which guarantees the authenticity and ownership of the audio data.

[2388] Step 8:

[2389] Download generated audio

[2390] Subject: Terminal (animation production company)

[2391] The device downloads the generated audio data and the original audio NFT from the server through an appropriate interface.

[2392] Step 9:

[2393] Anime production and revenue records

[2394] Subject: Terminal

[2395] The device creates and publishes an animated work using the downloaded generated voice. After publication, the device records the revenue from the work and transmits the data to a server.

[2396] Step 10:

[2397] Revenue sharing

[2398] Subject: Server

[2399] The server receives the revenue data and distributes the revenue to the voice actors based on a pre-defined percentage. The revenue distribution process is carried out via electronic payment.

[2400] Step 11:

[2401] Emotion Recognition and Analysis

[2402] Subject: User (viewer)

[2403] The user watches an anime work, and the emotion engine recognizes the emotion they feel. The user's emotion data is then sent to the server.

[2404] Subject: Server

[2405] The server receives and analyzes the emotion data sent by the user, and the analysis results are used to improve the generative AI model.

[2406] Step 12:

[2407] Gathering feedback

[2408] Subject: User (viewer)

[2409] Users provide feedback to the platform about the quality and impressions of the generated speech, and the feedback is sent to the server.

[2410] Step 13:

[2411] Feedback analysis and model refinement

[2412] Subject: Server

[2413] The server receives and analyzes feedback from viewers and emotion analysis data from the emotion engine. The analysis results are used to improve the generative AI model and increase the quality of the next generated voice.

[2414] Examples:

[2415] Suppose the voice actor playing main character A in a long-running anime series has to take a break due to illness. Past voice data from the voice actor is recorded on a device and uploaded to a server. The server analyzes it and trains a generative AI model to generate new voice. The generated voice is then evaluated for quality and assigned an NFT. The animation production company downloads the generated voice and produces and releases the work, with a portion of the profits going back to the voice actor. Viewers watch the work, and the emotion engine analyzes their emotions, sending feedback to the server. The server uses the analysis results to improve the AI ​​model and increase the quality of the next generated voice. This system simultaneously maintains the quality of the voice actor's voice while also improving it to match the viewer's emotions.

[2416] Example 2

[2417] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2418] In traditional anime production, voice recording by voice actors is essential, and missing voice recording directly affects the quality of the work. If a voice actor is unable to participate in recording due to illness or other reasons, the risk of production delays increases. Voice actors' earnings also tend to be unstable. Furthermore, it is difficult to generate voices that reflect the viewer's emotions, making it difficult to consistently provide high-quality voices. Our goal is to solve these issues and provide a system that can stably provide higher-quality anime works.

[2419] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2420] In this invention, the server includes a voice analysis means, a training means for the AI ​​model, a means for evaluating the generated voice, a means for assigning NFTs to the voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion engine means for recognizing and analyzing viewers' emotional data, and a testing means for evaluating the quality of the generated voice. This allows for the generation of high-quality voice even when the voice actor is unable to participate in recording, and enables the generation of voice that reflects the viewer's emotions while maintaining the quality of the work. Furthermore, by steadily returning revenue to the voice actor, economic stability can be achieved.

[2421] A "voice analysis means" is a means for recording a voice actor's voice data and breaking down and analyzing the voice into characteristics such as waveform, spectrum, pitch, volume, and intonation.

[2422] "Means for training a generative AI model" refers to means for training a generative AI model using analyzed voice feature data so that it can reproduce the unique texture and tone of a voice actor.

[2423] The "means for evaluating generated speech" is a means for evaluating whether the generated speech is of appropriate quality, and uses speech quality evaluation metrics.

[2424] "Method of assigning NFTs to audio data" is a method of generating NFTs (non-fungible tokens) for a voice actor's original audio data, registering the NFTs on the blockchain, and proving the ownership and authenticity of the audio data.

[2425] "Means for downloading generated audio" refers to the means by which animation production companies can download the generated audio data and NFTs to their own devices.

[2426] A "revenue distribution vehicle" is a means for recording revenue earned after the release of an anime work and returning it to voice actors based on a preset percentage.

[2427] The "means for collecting feedback from viewers" refers to a means for collecting and analyzing feedback and emotional data provided by viewers after watching an anime work.

[2428] The "emotion engine means" is a means for recognizing and analyzing the viewer's emotions, and the data is used to improve the generative AI model.

[2429] "Testing means for evaluating the quality of generated speech" refers to a means for conducting tests to evaluate the quality of speech generated by a trained AI model.

[2430] MODE FOR CARRYING OUT THE INVENTION

[2431] This invention is a system that uses generative AI technology to replicate a voice actor's voice and assigns an NFT to preserve the value of the original. Furthermore, this system combines an emotion engine that recognizes the user's emotions, analyzes the impact of the generated voice on the viewer, and can improve the AI ​​model based on that feedback. This system can maintain the quality of anime works and stabilize the voice actor's income even when the voice actor is unable to participate in recording.

[2432] Overall system configuration

[2433] The system of the present invention comprises the following elements:

[2434] 1. Voice analysis method: Analyze the voice actor's voice data.

[2435] 2. Training method for generative AI model: Train the generative AI model based on the analyzed data.

[2436] 3. Evaluation method for generated speech: Evaluate the quality of the generated speech.

[2437] 4. Method of assigning NFTs to audio data: NFTs will be assigned to the voice actor's original audio data and registered on the blockchain.

[2438] 5. Method for downloading generated audio: Animation production companies will download generated audio and NFTs.

[2439] 6. Revenue distribution method: Revenue from the anime works produced will be recorded and returned to the voice actors.

[2440] 7. A method for collecting viewer feedback: Collect and analyze viewer feedback to help improve the AI ​​model.

[2441] 8. Emotion engine means: Recognize and analyze viewer emotions and use that data to improve the generative AI model.

[2442] 9. Testing method to evaluate the quality of generated speech: Evaluate the quality of speech generated by the trained AI model.

[2443] Hardware and software used

[2444] Hardware

[2445] Recording device: A high-quality microphone (e.g., Shure SM7B), a PC for recording audio (e.g., a general-purpose laptop)

[2446] Server: A high-performance server in a data center (e.g., a virtual machine in a cloud environment)

[2447] Animation production device: PC for animation production (e.g., workstation equipped with an LCD pen tablet)

[2448] software

[2449] Voice analysis software: General-purpose voice analysis tools

[2450] Generative AI Models: A General-Purpose Deep Learning Framework

[2451] NFT Blockchain Platform: A Popular Distributed Ledger Technology

[2452] Electronic payment system: Online payment service

[2453] Feedback Collection Platform: Survey Tools

[2454] Explanation of system processing

[2455] The specific processing flow of this system will be explained below.

[2456] Audio data collection and analysis

[2457] Subject: Terminal

[2458] The device records the voice actor's voice using a high-quality microphone and generates an audio file.

[2459] The collected audio files contain the characters' lines and express various emotions.

[2460] The device uploads the audio file to the server.

[2461] Subject: Server

[2462] The server uses audio analysis software to analyze the received audio files.

[2463] The server breaks down the audio into features such as waveform, spectrum, pitch, volume, and intonation, and generates analysis data.

[2464] The server stores the analyzed data in a database.

[2465] Training generative AI models

[2466] Subject: Server

[2467] The server trains a generative AI model based on the analyzed voice feature data.

[2468] The generative AI model runs on a common deep learning framework.

[2469] Once trained, the generative AI model will be able to reproduce the unique texture and tone of a voice actor.

[2470] Evaluation and improvement of generated speech

[2471] Subject: Server

[2472] The server evaluates the quality of the generated voice to test new voices created by the generative AI model.

[2473] Voice quality assessment metrics are used in the testing.

[2474] Based on the test results, the parameters of the AI ​​model are readjusted and retrained.

[2475] NFT assignment for original audio

[2476] Subject: Server

[2477] The server generates an NFT for the voice actor's original voice.

[2478] The generated NFT is registered on the blockchain, proving ownership and authenticity of the audio data.

[2479] Audio Usage and Revenue Sharing

[2480] Subject: Terminal (animation production company)

[2481] The animation production company's terminal downloads the generated audio data and NFT from the server.

[2482] The downloaded audio data will be used in the production of anime works.

[2483] After the anime work is released, revenue data is uploaded from the terminal to the server.

[2484] Subject: Server

[2485] The server receives the revenue data and returns the revenue to the voice actor based on a preset percentage.

[2486] An electronic payment system will be used for revenue sharing.

[2487] Audience emotion recognition and feedback collection

[2488] Subject: User (viewer)

[2489] Users can view the completed animation and enter their impressions in a feedback form.

[2490] The emotion engine recognizes and collects audience emotional data.

[2491] Subject: Server

[2492] The server analyzes the feedback from viewers and the emotional data from the emotion engine.

[2493] The analysis results will be used to further improve the AI ​​model, contributing to improving the quality of generated speech.

[2494] Examples and prompts

[2495] For example, this system would be effective if a voice actor playing a key character in an anime had to take a break due to illness. The voice data of the voice actor recorded on the device is uploaded to a server and analyzed. A generative AI model is trained based on the analysis results, and a new voice is generated. The original voice is assigned an NFT and registered on the blockchain. The animation production company downloads the generated voice and completes the work. Revenue is returned to the voice actor, and viewer emotional data and feedback are used to improve the AI ​​model.

[2496] Example prompt sentence:

[2497] "Generate the following line in the voice of the specified voice actor. Line: 'Hello, this is Character A. It's a great day today.'"

[2498] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2499] Divide the processing flow of this system's program into processing steps

[2500] Step 1: Collect and upload audio data

[2501] Step 2: Analyzing the audio data

[2502] Step 3: Training the generative AI model

[2503] Step 4: Evaluate and improve the generated speech

[2504] Step 5: Adding an NFT to the original audio

[2505] Step 6: Download and use the generated audio

[2506] Step 7: Revenue sharing

[2507] Step 8: Recognize audience emotions and collect feedback

[2508] Description of each processing step

[2509] Step 1: Collect and upload audio data

[2510] Subject: Terminal

[2511] The device uses a high-quality microphone to record voice data from voice actors, including the character's lines and various emotional expressions.

[2512] The device temporarily stores the recorded audio data and prepares it for uploading to the server.

[2513] Upload the audio file to the server using a secure communication protocol (e.g. HTTPS).

[2514] Input: Recorded audio file

[2515] Output: Audio file uploaded to the server

[2516] Step 2: Analyzing the audio data

[2517] Subject: Server

[2518] The server launches voice analysis software to analyze the received voice data.

[2519] The audio data is decomposed into features such as waveform, spectrum, pitch, volume, and intonation, and these data are generated.

[2520] The server stores the generated feature data in a database.

[2521] Input: Uploaded audio file

[2522] Output: Audio feature data

[2523] Step 3: Training the generative AI model

[2524] Subject: Server

[2525] The server uses the analyzed audio feature data to train a generative AI model.

[2526] Optimize the parameters of AI models using deep learning frameworks (e.g., TensorFlow, PyTorch).

[2527] The training process requires a lot of data and computational resources to reproduce the unique texture of the voice actor.

[2528] Input: Audio feature data

[2529] Output: A trained generative AI model

[2530] Step 4: Evaluate and improve the generated speech

[2531] Subject: Server

[2532] The server uses a generative AI model to generate new speech and evaluate its quality.

[2533] We use voice quality evaluation metrics (e.g., PESQ, STOI) and improve the AI ​​model based on the evaluation results.

[2534] The server repeats this process to improve the accuracy of the model.

[2535] Input: A trained generative AI model

[2536] Output: Evaluated generated speech data

[2537] Step 5: Adding an NFT to the original audio

[2538] Subject: Server

[2539] The server generates an NFT (non-fungible token) for the voice actor's original voice data.

[2540] NFTs are created and registered on a blockchain platform (e.g., Ethereum).

[2541] The generated NFT proves ownership and authenticity of the audio data.

[2542] Input: Original audio data

[2543] Output: NFT registered on the blockchain

[2544] Step 6: Download and use the generated audio

[2545] Subject: Terminal (animation production company)

[2546] The animation production company's terminal downloads the generated audio data and NFT from the server.

[2547] The device uses the downloaded audio data and applies it to the animation work being produced.

[2548] Input: Generated audio data, NFT

[2549] Output: Audio data and NFTs used as part of the anime production

[2550] Step 7: Revenue sharing

[2551] Subject: Terminal (animation production company)

[2552] The device records revenue data for released anime works.

[2553] Periodically upload revenue data to a server.

[2554] Subject: Server

[2555] The server receives the revenue data and distributes the revenue to the voice actors based on pre-set percentages.

[2556] Revenue sharing will be via electronic payment systems (e.g., online payment services).

[2557] Input: Revenue Data

[2558] Output: Revenue share to voice actors

[2559] Step 8: Recognize audience emotions and collect feedback

[2560] Subject: User (viewer)

[2561] Users can view the completed animation and enter their impressions in a feedback form.

[2562] The emotion engine recognizes and collects audience emotional data.

[2563] Input: Viewer sentiment data and feedback

[2564] Output: Emotion data and feedback collected on the server

[2565] Subject: Server

[2566] The server analyzes the feedback from viewers and the emotional data collected by the emotion engine.

[2567] The analysis results are used to improve the AI ​​model, leading to improved quality of the generated voice.

[2568] Input: Collected emotion data and feedback

[2569] Output: An improved generative AI model

[2570] (Application example 2)

[2571] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2572] In anime production, if a key voice actor is temporarily unavailable, it is difficult to complete the work while maintaining the quality of that character's voice. Furthermore, while viewer feedback is important for improving the quality of voice data generated by generative AI models, traditional feedback collection methods have the problem of being unable to recognize user emotions in detail. Furthermore, there is a need for a mechanism to guarantee the authenticity and ownership of generated voices and to appropriately distribute revenue.

[2573] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice analysis means, a means for training the generative AI model, a means for evaluating the generated voice, a means for assigning NFTs to voice data, a means for downloading the generated voice, a revenue distribution means, a means for collecting feedback from viewers, an emotion recognition means, and a means for analyzing the collected emotion data and improving the generative AI model based on the analysis results. This makes it possible to generate high-quality voice even when the main voice actor is unavailable, collect feedback based on viewer emotions, improve the generative AI model, and improve the quality of the generated voice. Furthermore, assigning NFTs to the generated voice guarantees the authenticity and ownership of the voice data and enables appropriate revenue distribution.

[2574] "Voice analysis means" refers to a means for analyzing the voice data of a voice actor and extracting its characteristics.

[2575] "Means for training a generative AI model" means means for training a generative AI model based on analyzed audio data.

[2576] "Means for evaluating generated voice" refers to a means for evaluating the quality of voice generated by the generation AI.

[2577] "Means for assigning NFTs to audio data" refers to a means for assigning NFTs to generated audio data and registering it on the blockchain.

[2578] The "means for downloading generated voice" is a means for downloading generated voice data.

[2579] "Revenue distribution means" means a means for recording and managing revenues from generated voices and animation works, and distributing revenues to voice actors and related parties.

[2580] The "means for collecting feedback from viewers" is a means for collecting feedback from viewers and analyzing the data.

[2581] An "emotion recognition means" is a means for recognizing the viewer's emotions and collecting data on them.

[2582] "Means for analyzing collected emotional data and improving the generative AI model based on the results of the analysis" refers to means for analyzing emotional data collected from viewers and improving the generative AI model based on the results of the analysis.

[2583] The system of the present invention is realized by combining various means as follows. Specific embodiments are shown below.

[2584] 1. Collection and analysis of audio data

[2585] Subject: Terminal

[2586] Voice actors' voice data is recorded using high-quality recording equipment, and this voice includes the characters' lines and various emotional expressions.

[2587] The recorded audio data is uploaded from the device to a server, using high security and data compression technology to prevent data loss.

[2588] Subject: Server

[2589] The server receives the uploaded audio data and analyzes characteristics such as the audio waveform, spectrum, pitch, and tempo. This analysis is performed using libraries such as "librosa."

[2590] The analyzed audio feature data is stored in a database for use in training subsequent generative AI models.

[2591] 2. Training a generative AI model

[2592] Subject: Server

[2593] The server trains a generative AI model based on the speech feature data for analysis, using machine learning frameworks such as the "transformers" library.

[2594] The model is tuned to reproduce the unique texture of a particular voice actor's voice.

[2595] 3. Evaluation of generated speech

[2596] Subject: Server

[2597] The server generates new audio using the trained generative AI model and evaluates its quality, applying criteria for voice quality assessment and further fine-tuning the model if necessary.

[2598] 4. NFT assignment to original audio

[2599] Subject: Server

[2600] The server assigns an NFT to the voice actor's original voice data and registers the NFT on the blockchain, using an external blockchain API.

[2601] A registered NFT serves as proof of ownership and authenticity of the audio data.

[2602] 5. Audio Use and Revenue Sharing

[2603] Subject: Terminal (animation production company)

[2604] The animation production company downloads the generated audio data and NFTs from the server using a dedicated interface.

[2605] The downloaded generated audio is used to create an animated work, and after the work is released, revenue data is recorded and sent to a server.

[2606] Subject: Server

[2607] The server returns the revenue to the voice actor based on the received revenue data at a preset rate, using electronic payment technology.

[2608] 6. Audience Emotion Recognition and Feedback Collection

[2609] Subject: User (viewer)

[2610] Users can watch the finished animation and provide feedback on the quality of the generated voice, while the system recognizes the user's emotions in real time using the "transformers" library.

[2611] Subject: Server

[2612] The server collects and analyzes viewer feedback and emotion recognition results, which are used to improve the next generation AI model.

[2613] Specific examples

[2614] For example, there may be an anime featuring a character played by a well-known voice actor, but that voice actor is suddenly unable to appear. In this case, previously recorded voice data expressing the actor's lines and emotions is uploaded to a server, and a generative AI model is trained based on that data. Furthermore, by entering a prompt sentence as follows, character voices for specific situations can be generated.

[2615] (Example prompt): "Good morning, let's do our best today!" (in a cheerful tone)

[2616] This voice is used in the animation, and viewers can provide emotional feedback, such as "I'm excited." This feedback is analyzed in detail using emotion recognition technology and used to further impr...

Claims

1. A means of voice analysis; A means of training the generative AI model; A means for evaluating the generated speech; A means for assigning NFT to audio data; A means for downloading the generated audio; Revenue sharing instruments; A means of collecting feedback from viewers; A system including:

2. 10. The system of claim 1, further comprising a voice analysis means for analyzing voice data of a voice actor and training a generative AI model based on the analyzed data.

3. The system of claim 1, further comprising an NFT assignment means for assigning an NFT to the voice actor's original voice data and registering the NFT on a blockchain.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A