System
A system enables voice actors to register and manage their voice data, allowing filmmakers to efficiently search and convert text to speech, addressing the challenges of voice integration in anime production, enhancing production efficiency and user control.
Patent Information
- Application Number
- JP2024124066
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Finding suitable voice actors and incorporating their voices into anime or test footage is difficult for amateur producers and small production companies, requiring direct contact and auditions, which is time-consuming and costly, and limits voice actors' self-promotion and control over their voice usage.
A system allowing voice actors to register their voice data and usage conditions, enabling filmmakers to search for and convert input text into speech using voice generation AI, and incorporate the generated speech data into videos.
Facilitates efficient production by allowing amateur producers to easily use appropriate voice actors' voices, improving production efficiency and providing a platform that satisfies both parties.
Smart Images

Figure 2026022549000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When producing original anime or test footage, finding suitable voice actors and incorporating their voices into the footage can be extremely difficult, especially for amateur producers and small production companies. Traditional methods require direct contact with voice actors and auditions, which requires a great deal of effort and expense, lengthening the production period. Furthermore, voice actors have limited opportunities for self-promotion and casting, making it difficult to properly control how their voices are used. This leaves the needs of both parties unmet. [Means for solving the problem]
[0005] The system of the present invention includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into speech based on the voice data of the selected voice actor, and a means for incorporating the generated speech data into video. Using this system, even amateur producers and small production companies can easily incorporate appropriate voice actors' voices into their works, and voice actors can also clearly set the conditions for using their own voices. This improves the efficiency of production work and provides an environment that satisfies both voice actors and producers.
[0006] "Filmmaker" refers to a person or group that produces their own animated works or test footage.
[0007] "Voice data" refers to digital data of recorded voices provided by voice actors or voices generated by voice generation means.
[0008] A "voice actor" is a person who provides their voice and registers it by setting the terms and fees for its use.
[0009] "Terms of Use" refers to the rules and restrictions that govern how a voice actor may use their voice data.
[0010] A "means" refers to a method, device, or process employed to accomplish a particular purpose.
[0011] "Keywords" refers to the strings or words that filmmakers use to search for voice actors.
[0012] "Voice generation AI" refers to artificial intelligence technology that generates voice data based on input text.
[0013] "Database" refers to an electronic recording device or system for organizing and storing information.
[0014] "Text" refers to the character string or sentence that is input when generating speech.
[0015] "System" refers to a series of processing systems and services that are configured by combining the above means, data, devices, etc.
[0016] "Convert to speech" refers to the process of converting input text data into speech data.
[0017] "Incorporating" refers to the act of integrating the generated audio data into the video data to complete it as a single work. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention relates to a system for easily generating audio data and incorporating it into a self-produced animation or test video. Specific embodiments of the system of the present invention will be described below.
[0040] System Overview
[0041] This system consists of three main components.
[0042] 1. A means for voice actors to register their own voice data and terms of use
[0043] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[0044] 3. Means for converting input text into speech based on the voice data of the selected voice actor and incorporating the generated speech data into the video.
[0045] Program processing
[0046] 1. Voice actor voice data registration
[0047] A user (voice actor) accesses the website and creates an account.
[0048] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[0049] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[0050] The server validates the uploaded audio data and stores it in a database.
[0051] 2. Search and select a voice actor
[0052] The user (video creator) logs in to the website and accesses the voice actor search page.
[0053] The user (video creator) enters keywords and performs a search.
[0054] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[0055] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0056] 3. Generating audio data
[0057] Enter text on the detailed information page of the voice actor selected by the user (video producer).
[0058] The server uses speech generation AI to convert the input text into audio data.
[0059] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[0060] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[0061] Specific examples
[0062] For example, consider the case where a user (filmmaker) is looking for a "young female voice" to use in one of their own animation works. The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the results. The user selects an appropriate voice actor from the results and enters the text "Hello, nice to meet you" on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed easily.
[0063] In this way, the system of the present invention provides an environment in which users can quickly and efficiently generate audio data and use it in video productions, significantly reducing production costs and time and enabling even amateur producers to achieve professional-looking results.
[0064] The processing flow will be explained below.
[0065] Voice actor voice data registration process
[0066] Step 1:
[0067] The user (voice actor) accesses the website and opens the account creation page.
[0068] The device presents the user with an account creation form.
[0069] Step 2:
[0070] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[0071] The terminal sends the input information to the server.
[0072] Step 3:
[0073] The server creates a new user record in the database based on the received account information.
[0074] The server validates the input and saves it to the database.
[0075] Step 4:
[0076] The server sends a notification to the user that account registration is complete.
[0077] The server sends emails and site notifications.
[0078] Step 5:
[0079] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[0080] The device displays a voice registration form.
[0081] Step 6:
[0082] The user (voice actor) fills out the required information in the form and uploads the audio file.
[0083] The device sends voice data and other input information to the server.
[0084] Step 7:
[0085] The server validates the uploaded audio data and stores it in the database.
[0086] The server checks the format and size of the audio data and stores it in a database.
[0087] Step 8:
[0088] The server sends a notification to the user that the voice data has been registered.
[0089] The server sends emails and site notifications.
[0090] Voice Actor Search and Selection Process
[0091] Step 1:
[0092] The user (video creator) logs in to the website and accesses the voice actor search page.
[0093] The terminal displays a login form and, after successful authentication, a search page.
[0094] Step 2:
[0095] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[0096] The terminal sends the entered keyword to the server.
[0097] Step 3:
[0098] The server searches the database for voice actors that match the keywords.
[0099] The server executes a search query against the database and retrieves the results.
[0100] Step 4:
[0101] The server sends the search results to the user's (filmmaker's) device.
[0102] The server sends the search results in JSON format, which the device parses for display.
[0103] Step 5:
[0104] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0105] The device will display a detailed information page.
[0106] Voice data generation process
[0107] Step 1:
[0108] On the detailed information page of the voice actor selected by the user (video producer), enter a line (e.g., "Hello, nice to meet you") in the text input box and click the "Generate Voice" button.
[0109] The device sends the entered text and the selected voice actor's ID to the server.
[0110] Step 2:
[0111] The server sends a voice generation request to the voice generation AI.
[0112] The server sends the text and voice actor ID to the voice generation API and receives the voice data.
[0113] Step 3:
[0114] The server converts the generated audio data into a specific format and sends it back to the user (video producer).
[0115] The server converts the audio data into MP3 or WAV format and sends it to the device.
[0116] Step 4:
[0117] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[0118] The device displays a download link for the audio data, which the user can then download and use.
[0119] Specific examples
[0120] For example, a user (video creator) searches for "the voice of a young woman," selects a specific voice actor, enters the text "Hello, nice to meet you," and generates the voice. Through this process, a series of steps are realized in which the generated voice data is received from the server, downloaded, and incorporated into the user's own animated work.
[0121] Example 1
[0122] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0123] In today's video production, generating and incorporating audio data requires a significant amount of time and cost. Searching, selecting, and managing audio from professional voice actors is particularly complex, and the lack of an efficient system is a problem. This makes amateur productions and short-term projects difficult.
[0124] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0125] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for incorporating the generated voice data into a video work, a means for converting input text into voice using a voice generation AI model, and a means for providing the generated voice data to users, thereby enabling video producers to efficiently generate and incorporate voice data.
[0126] A "voice actor" is an individual or group that uses their own voice to provide audio data.
[0127] "Audio data" refers to audio information provided by voice actors stored in digital data format.
[0128] "Conditions of use" refers to restrictions and permissions regarding the use of audio data, such as whether commercial use is permitted or not, and the expiration date of use.
[0129] "Filmmaker" means an individual or organization that produces a film work.
[0130] A "keyword" is a character string or phrase that a user enters to search for a voice actor.
[0131] A "voice generation AI model" is a type of artificial intelligence technology that analyzes input text and generates it as voice data.
[0132] "Text" refers to the written information that is the basis for generating speech.
[0133] A "server" is a computer system that processes various types of data and provides services to users.
[0134] A "database" is a system that systematically stores and manages information such as voice data and usage conditions.
[0135] "Prompt sentence" refers to an example of text that a user inputs into a speech-generation AI model.
[0136] This invention relates to a system that allows filmmakers to efficiently generate audio data and incorporate it into their video works. This system has the means to register, search, generate, and provide audio data, and by using a generative AI model, it realizes rapid and high-quality audio data generation.
[0137] System Program Overview
[0138] The program of this system is implemented using the following hardware and software.
[0139] Server: A computer system for storing, searching, generating, and providing data.
[0140] Database: Management of voice data, terms of use, user information, etc.
[0141] Website: The interface for users (voice actors and filmmakers) to access
[0142] Speech generation AI model: Artificial intelligence technology that analyzes input text and generates voice data
[0143] Voice data registration details
[0144] A user (voice actor) accesses the website and creates an account. When creating an account, the user enters the required information and clicks "Register," and the server saves the information in a database. Once registration is complete, the user (voice actor) logs in and registers their voice data, terms of use, and fees. Once the voice data is uploaded, the server validates it and saves it in a database.
[0145] More on finding and selecting voice actors
[0146] The user (filmmaker) logs in to the website and accesses the voice actor search page. Using the search function, they enter a keyword, such as "female female voice," and execute a search. The server searches the database for the relevant voice actor and presents the results to the user (filmmaker). The user can then check the detailed information from the displayed list of voice actors and select the one they need.
[0147] Voice data generation details
[0148] The user (video producer) enters text, such as lines, on the detailed information page of the voice actor they selected. When they click the "Generate" button, the server uses a voice generation AI model to convert the entered text into audio data. The generated audio data is then converted by the server into an appropriate format and provided to the user. The user (video producer) can then download this audio data and incorporate it into their own video work.
[0149] Specific examples
[0150] For example, if a user (filmmaker) is looking for a "young female voice" to use in their own animation work, they can proceed as follows: The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the search results. The user selects an appropriate voice actor from the list and enters text, such as "Hello, nice to meet you," on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed.
[0151] Example prompt sentence:
[0152] "Please convert the text "Hello, nice to meet you" into speech data in a young female voice."
[0153] In this way, the system of the present invention is a mechanism that realizes efficient and high-quality audio data generation and provides video producers with audio data that can be used quickly.
[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0155] Step 1:
[0156] User (voice actor) account creation and information registration
[0157] The user (voice actor) accesses the website and enters the necessary information into the "Create Account" form, such as name, email address, and password.
[0158] When a user clicks "Register," the server receives the information and performs validation, such as checking the format of the email address and the strength of the password.
[0159] If validation passes, the server saves the account information in the database, allowing the user to log in.
[0160] After logging in, the user (voice actor) registers their own voice data, terms of use, and usage fees. After uploading the voice data, the server validates the data and stores it in the database. Specifically, it checks the format and size of the voice file.
[0161] Input: Account information entered by the user (name, email address, password) and voice data
[0162] Output: Account information and voice data stored in a database
[0163] Step 2:
[0164] User (video creator) login and voice actor search
[0165] The user (video creator) logs in to the website and accesses the voice actor search page. A username and password are required to log in.
[0166] The user (filmmaker) enters a keyword into the search bar, for example, "young woman's voice," and executes the search.
[0167] The server searches the database for a list of matching voice actors and retrieves the results, including the voice data and usage conditions that match the search criteria.
[0168] The server displays the search results to the user (video creator), which include the name of the voice actor, sample audio, terms of use, etc.
[0169] Input: The keyword that the user types into the search bar
[0170] Output: A list of voice actors displayed as search results
[0171] Step 3:
[0172] User (video creator) voice actor selection and text input
[0173] The user (filmmaker) selects the appropriate voice actor from the search results and accesses their details page.
[0174] The user checks the details and enters the text they want to convert into speech in the text entry field, for example, "Hello, nice to meet you."
[0175] The user clicks the "Generate" button, which sends the entered text to the server.
[0176] Input: The voice actor the user selects, and the text they enter into the text field.
[0177] Output: Text data sent to the server
[0178] Step 4:
[0179] Voice data generation by the server
[0180] The server sends the input text to a speech generation AI model, which analyzes the text and generates appropriate speech data.
[0181] The voice generation AI model generates speech based on the characteristics of a specified voice actor, for example, generating "Hello, nice to meet you" in a young female voice.
[0182] The server receives the generated audio data and performs format conversion and quality checks, including encoding the audio data and removing noise.
[0183] Input: Text data sent to the server
[0184] Output: Generated audio data
[0185] Step 5:
[0186] Providing audio data to users (video producers)
[0187] The server generates a download link for the generated audio data and provides the link to the user (video producer).
[0188] The user (video creator) downloads the audio data from the provided link and can incorporate the downloaded audio data into their own video work.
[0189] Input: Generated audio data
[0190] Output: Download link and audio data provided to the user
[0191] (Application example 1)
[0192] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0193] In existing content distribution systems, generating audio data and incorporating it in real time is extremely time-consuming, especially when providing an interactive experience. Furthermore, there is a lack of efficient means for audiovisual producers to search for and use the specific audio data they require. This creates a need for technology that can quickly and efficiently generate high-quality audio data and provide interactive experiences through visual display devices in real time.
[0194] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0195] In this invention, the server includes means for audio providers to register their own audio data and usage conditions, means for content creators to search for audio providers using keywords, means for converting input text into audio based on the audio data of the selected audio provider, means for incorporating the generated audio data into content in real time, and means for providing an interactive experience through a visual display device, thereby making it possible to quickly and efficiently generate audio data and provide an interactive experience through a visual display device in real time.
[0196] "Audio Data" means a digital file of sound recorded or generated by an audio provider.
[0197] "Audio provider" is a person or organization that provides audio data and sets the terms of use.
[0198] "Conditions of Use" are specific conditions and restrictions set by the audio provider when using audio data.
[0199] A "content creator" is a person or organization that creates content such as video or audio.
[0200] "Keywords" are specific words or phrases used in searches.
[0201] A "generative AI model" is an artificial intelligence technology that takes input text and converts it into voice data.
[0202] "Real time" refers to a state in which processing and responses occur almost simultaneously.
[0203] A "visual display device" is a device that allows a user to visually view content, and includes smartphones and head-mounted displays.
[0204] An "interactive experience" is one in which a user can interact with a system and receive responses in real time.
[0205] "Content" refers to digital media that includes information such as video and audio.
[0206] The present invention provides a system for generating audio data and incorporating it into content in real time, and is specifically implemented with the following configuration.
[0207] 1. System Overview
[0208] The system includes a voice provider, a content creator, a server, and a visual display device. The voice provider registers their voice data and terms of use, and the content creator searches for the voice provider using keywords, inputs text, and uses a voice generation AI model to generate voice data, which is then incorporated into the content in real time.
[0209] 2. Hardware and Software
[0210] Hardware: Smartphones, head-mounted displays, servers
[0211] Software: Web-based applications, databases (e.g., MySQL), speech generation AI models (e.g., Tacotron, WaveNet)
[0212] 3. System Processing Procedures
[0213] 1. Voice provider voice data registration
[0214] The voice provider visits the website and creates an account.
[0215] Once you enter your account information and register, the server stores that information in a database.
[0216] After completing the registration, the voice provider inputs his / her voice data, the terms of use, and the usage fee, and transmits them to the server.
[0217] The server validates the uploaded audio data and stores it in a database.
[0218] 2. Search and select an audio provider
[0219] A content creator logs into a website and visits the search page.
[0220] A content creator enters keywords and performs a search.
[0221] The server searches the database for the appropriate audio provider and presents the results to the content creator.
[0222] The content creator checks the detailed information from the displayed list of audio providers and selects one.
[0223] 3. Generating audio data
[0224] Enter text on the details page of the audio provider that the content creator has selected.
[0225] The server uses a generative AI model to convert the input text into audio data.
[0226] After successful speech generation, the server converts the generated speech data into an appropriate format and provides an interactive experience through a visual display device in real time.
[0227] 4. Specific Examples
[0228] For example, if a content creator searches for "young female voice" and enters the text "Hello, nice to meet you," the server will use the generative AI model to convert this text into speech data and incorporate the generated speech data into the content in real time, thereby providing an interactive experience to the user through a visual display device.
[0229] 5. Examples of prompts
[0230] Generative AI models:
[0231] Prompt: Convert the following text into the voice of the voice actor you provided: "Hello, nice to meet you."
[0232] Expected output: The text "Hello, nice to meet you" will be output as audio data from the selected young female voice actress.
[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0234] Step 1:
[0235] The voice provider accesses the website and creates an account. The account information entered by the voice provider is sent to the server, which stores the information in a database.
[0236] Input: Account information input by the voice provider
[0237] Data processing: The server receives the account information and stores it in the database in the appropriate format.
[0238] Output: Account registration completion notification
[0239] Step 2:
[0240] After completing registration, the voice provider enters their voice data, terms of use, and usage fees into a form on the website and submits it to the server, which then validates the uploaded voice data and stores it in a database.
[0241] Input: Input of audio data, terms of use, and fees by audio provider
[0242] Data processing: The server receives the data, validates the voice data, and saves it to the database.
[0243] Output: Notification of completion of voice data registration
[0244] Step 3:
[0245] A content creator logs in to the website and accesses the search page. The content creator enters keywords and performs a search. The server searches the database for the corresponding audio provider and presents the results to the content creator.
[0246] Input: Keywords entered by content creators
[0247] Data processing: The server performs a keyword search and extracts the corresponding voice providers from the database.
[0248] Output: Presenting search results
[0249] Step 4:
[0250] The content creator checks the details of the audio providers from the list displayed and selects one. On the details page of the selected audio provider, they input text and request audio generation.
[0251] Input: Content creator selects audio provider and inputs text
[0252] Data processing: The server receives the text and invokes a generative AI model based on the voice data of the selected voice provider.
[0253] Output: Notification of execution of generation request
[0254] Step 5:
[0255] The server uses the generative AI model to convert the input text into speech data, and after successful speech generation, converts the generated speech data into an appropriate format and incorporates it into the content in real time.
[0256] Input: Text input to the generative AI model
[0257] Data processing: The generative AI model converts the text into audio data, and the server retrieves the audio data.
[0258] Output: Generated audio data
[0259] Step 6:
[0260] The generated audio data is incorporated into the content through a visual display device to provide an interactive experience for the user.
[0261] Input: Generated audio data
[0262] Data processing: Integrating audio data into the content of a visual display
[0263] Output: Interactive experience through a visual display device
[0264] This series of processing steps allows content creators to efficiently generate audio data and deliver real-time interactive experiences through visual display devices.
[0265] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0266] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user-created animation or test video. A specific embodiment of the system of the present invention will be described below.
[0267] System Overview
[0268] This system consists of four main components.
[0269] 1. A means for voice actors to register their own voice data and terms of use
[0270] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[0271] 3. Means for converting input text into speech based on the voice data of the selected voice actor.
[0272] 4. A means to recognize the user's emotions using an emotion engine and generate voice data based on that information
[0273] Program processing
[0274] 1. Voice actor voice data registration
[0275] A user (voice actor) accesses the website and creates an account.
[0276] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[0277] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[0278] The server validates the uploaded audio data and stores it in a database.
[0279] 2. Search and select a voice actor
[0280] The user (video creator) logs in to the website and accesses the voice actor search page.
[0281] The user (video creator) enters keywords and performs a search.
[0282] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[0283] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0284] 3. Emotion-Recognition Assistance
[0285] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[0286] The device recognizes the user's emotions in real time from their facial expressions and voice, and sends that information to the server.
[0287] The emotion engine assists in generating voice data containing appropriate intonation and emotional expression based on the recognized emotion.
[0288] 4. Generating Audio Data
[0289] The server uses a speech generation AI and emotion engine to convert the input text into voice data.
[0290] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[0291] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[0292] Specific examples
[0293] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. In this process, the emotion "joy" is recognized from the user's facial expressions and voice, and the server selects voice data from the voice actor based on that, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[0294] In this way, the system of the present invention can significantly improve the quality of video works by recognizing the user's emotions in real time and providing appropriate audio data, thereby significantly reducing production costs and time and providing an easy-to-use environment for users to easily create high-quality works.
[0295] The processing flow will be explained below.
[0296] Voice actor voice data registration process
[0297] Step 1:
[0298] The user (voice actor) accesses the website and opens the account creation page.
[0299] The device presents the user with an account creation form.
[0300] Step 2:
[0301] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[0302] The terminal sends the input information to the server.
[0303] Step 3:
[0304] The server creates a new user record in the database based on the received account information.
[0305] The server validates the input and saves it to the database.
[0306] Step 4:
[0307] The server sends a notification to the user that account registration is complete.
[0308] The server sends emails and site notifications.
[0309] Step 5:
[0310] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[0311] The device displays a voice registration form.
[0312] Step 6:
[0313] The user (voice actor) fills out the required information in the form and uploads the audio file.
[0314] The device sends voice data and other input information to the server.
[0315] Step 7:
[0316] The server validates the uploaded audio data and stores it in the database.
[0317] The server checks the format and size of the audio data and stores it in a database.
[0318] Step 8:
[0319] The server sends a notification to the user that the voice data has been registered.
[0320] The server sends emails and site notifications.
[0321] Voice Actor Search and Selection Process
[0322] Step 1:
[0323] The user (video creator) logs in to the website and accesses the voice actor search page.
[0324] The terminal displays a login form and, after successful authentication, a search page.
[0325] Step 2:
[0326] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[0327] The terminal sends the entered keyword to the server.
[0328] Step 3:
[0329] The server searches the database for voice actors that match the keywords.
[0330] The server executes a search query against the database and retrieves the results.
[0331] Step 4:
[0332] The server sends the search results to the user's (filmmaker's) device.
[0333] The server sends the search results in JSON format, which the device parses for display.
[0334] Step 5:
[0335] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0336] The device will display a detailed information page.
[0337] Assisted processes using emotion recognition
[0338] Step 1:
[0339] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[0340] The terminal displays an input text box.
[0341] Step 2:
[0342] The user (video creator) uses a camera and microphone to provide facial expressions and voice to the system.
[0343] The device captures the user's facial expressions and voice data in real time and sends it to the server.
[0344] Step 3:
[0345] The server uses an emotion engine to recognize the user's emotions in real time from the transmitted facial and voice data.
[0346] The server sends the input data to the emotion engine and analyzes the emotions.
[0347] Step 4:
[0348] Based on the recognized emotion, the server sends a request to the speech generation AI to generate speech data containing expressions that best fit the input text.
[0349] The server sends the emotion information and text to the speech generation AI.
[0350] Step 5:
[0351] The server receives the generated audio data, converts it into an appropriate format, and provides it to the user (video producer).
[0352] The server converts the audio data into MP3 or WAV format and sends it to the device.
[0353] Step 6:
[0354] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[0355] The device displays a download link for the audio data, which the user can then download and use.
[0356] Specific examples
[0357] For example, a user (video producer) searches for "a young female voice," selects a specific voice actor, and enters the text "Hello, nice to meet you." Next, the user uses a camera or microphone to provide the system with their own smiling face and cheerful voice, which the emotion engine recognizes as "joy." Based on this emotional information, the server sends a request to the voice generation AI, which generates voice data reflecting the emotion of "joy." The user can download this generated voice data and incorporate it into their own animated work, enabling them to create an emotionally rich production.
[0358] Example 2
[0359] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0360] Conventionally, when generating audio data in video production, it has been difficult to provide audio that reflects the user's emotions. Furthermore, searching and managing voice data for voice actors is cumbersome, increasing production costs and time. This has led to the problem of making it difficult to efficiently produce high-quality video works.
[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0362] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for recognizing the user's emotions and generating voice data based on that information, and a means for incorporating the generated voice data into video. This makes it possible to efficiently generate high-quality voice data that reflects the user's emotions and incorporate it into video.
[0363] "Voice actors" are people who register their own voice data and play the role of providing audio for video production.
[0364] "Video producers" are people who generate the necessary audio data and perform editing work when producing a video work.
[0365] "Audio Data" means audio recordings provided by voice actors or audio generated by a generative AI model.
[0366] "Terms of Use" refers to the conditions and restrictions for using the voice data provided by the voice actors.
[0367] "Keywords" are words or phrases that filmmakers use to search for voice actors.
[0368] "Emotion recognition" is a technology that analyzes a user's facial expressions and voice and identifies their emotions in real time.
[0369] "Speech generation artificial intelligence" is a technology that generates natural-sounding speech based on input text and emotional data.
[0370] The "database" is an information collection system for efficiently managing and storing voice actor audio data and usage conditions.
[0371] "Text" refers to the string of characters entered by the video producer, and is the words or sentences that form the basis of the audio data.
[0372] "Video embedding" refers to the process of embedding the generated audio data as part of a video production.
[0373] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user's own animation or video work. A specific embodiment of the system of the present invention is described below.
[0374] This system consists of four main components.
[0375] 1. How to register voice data
[0376] A user (voice actor) accesses the website and creates an account by entering information such as name, email address, and password, and submitting it.
[0377] The server validates the entered information and saves it in the database. After successful saving, the user (voice actor) enters the voice data, terms of use, and fees, and uploads it.
[0378] The server validates the audio data and stores it in the database.
[0379] 2. How to search for voice actors
[0380] The user (video creator) logs in to the system and accesses the voice actor search page, where they enter keywords and perform a search.
[0381] The server searches the database for the relevant voice actor and displays the results as a list to the user (video producer).
[0382] The user (video creator) clicks and selects detailed information from the list of voice actors displayed.
[0383] 3. Text Input and Emotion Recognition
[0384] The user (video producer) goes to the details page of the voice actor they have chosen and enters the lines in the text input box.
[0385] The user's facial expressions and voice are captured in real time, and the device recognizes their emotions using emotion recognition software (e.g., OpenFace or Microsoft Emotion API).
[0386] The recognized emotion data is sent to a server.
[0387] 4. Audio data generation means
[0388] The server uses a speech generation AI (for example, Google Text-to-Speech API or Amazon Polly) to generate speech based on the input text and emotion data.
[0389] The server converts the generated audio data into an appropriate format and provides it to the user (video producer).
[0390] The user (video creator) downloads the generated audio data and incorporates it into their own video work.
[0391] Specific examples
[0392] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. During this process, the emotion "joy" is recognized from the user's facial expressions and voice. Based on this, the server selects voice data from the voice actor, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[0393] An example of a prompt to input to a generative AI model is as follows:
[0394] 1. "In a young female voice, please vocalize the lines "Hello" and "Nice to meet you" with a joyful emotion."
[0395] 2. "Can you say 'I'm looking forward to tomorrow!' in an energetic male voice with a surprised expression?"
[0396] 3. "Generate a recording of a middle-aged man quietly saying "goodbye" with sadness in his voice."
[0397] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[0398] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0399] Step 1: Create an account (a method for registering audio data)
[0400] Input: The user (voice actor) visits the website and enters their account information, including their name, email address, and password.
[0401] Data Processing: The server validates the entered information to ensure accuracy and format.
[0402] Output: If validation is successful, the server saves the account information to the database and sends the user a confirmation email.
[0403] Specific operation: The user (voice actor) enters the required information into the web form and clicks the "Create Account" button. The server saves the information in the database and automatically sends a confirmation email.
[0404] Step 2: Upload the audio data and terms of use
[0405] Input: After logging in, the user (voice actor) enters the audio data, terms of use, and usage fee into the web form and presses the upload button.
[0406] Data processing: The server receives the audio file and validates the file format and quality. It also checks the terms of use and fees.
[0407] Output: The voice data and terms of use that pass validation are saved in the database, and a success notification is displayed to the user.
[0408] Specific operation: The user (voice actor) selects an audio file, enters the required conditions, and clicks the "Upload" button. The server processes this and notifies the user of the result.
[0409] Step 3: Find a voice actor
[0410] Input: The user (video producer) logs into the system, accesses the voice actor search page, enters keywords, and presses the search button.
[0411] Data calculation: The server searches the database for the appropriate voice actor and filters the results that match the keywords.
[0412] Output: The search results are displayed as a list, and detailed information about the corresponding voice actors is presented to the user.
[0413] Specific operation: The user (filmmaker) enters keywords in the search box and clicks the "Search" button. The server retrieves the results and displays them on the page.
[0414] Step 4: Select a voice actor and enter text
[0415] Input: The user (video creator) clicks on the displayed voice actor details page and enters lines in the text input box.
[0416] Data processing: The server records the user's selection and temporarily stores the text data.
[0417] Output: A web page displays text entry boxes and other details that the user types and are sent to the server and stored.
[0418] Specific operation: The user (filmmaker) opens the details page, enters the dialogue text, and clicks the "Submit" button. The server receives and stores this information.
[0419] Step 5: Emotion Recognition
[0420] Input: While the user (video creator) is entering text, the device captures the user's facial expressions and voice in real time.
[0421] Data calculation: The device uses emotion recognition software (e.g., OpenFace or Microsoft Emotion API) to analyze the captured data and generate emotion data.
[0422] Output: The generated emotion data is sent to the server and associated with the text data.
[0423] How it works: While the user (filmmaker) is typing text, the device captures facial expressions and voice using the built-in camera and microphone, and runs emotion recognition software.
[0424] Step 6: Generate audio data
[0425] Input: The server sends the input text and emotion data to a speech generation AI (e.g., Google Text-to-Speech API or Amazon Polly).
[0426] Data calculation: The voice generation AI processes the input data based on the prompt sentence and generates voice data that reflects the emotion.
[0427] Output: The generated audio data is returned to the server and converted into the appropriate format.
[0428] Specific operation: The server passes text and emotion data to the speech generation AI, receives the generated speech data, and performs format conversion.
[0429] Step 7: Provide audio data
[0430] Input: The server provides the generated audio data to the user (video producer).
[0431] Data processing: The server rechecks the quality of the audio data and generates a download link.
[0432] Output: A page is displayed containing a link that allows the user (the filmmaker) to download the audio data.
[0433] Specific operation: The user (video creator) clicks on the download link, obtains the audio data, and incorporates it into the video work.
[0434] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[0435] (Application example 2)
[0436] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0437] A major challenge in current video production is the significant time and expense required to generate appropriate audio data. Particularly in the advertising field, emotive audio resonates with viewers and increases the effectiveness of advertising. However, existing speech synthesis technologies have limitations in expressing emotions, making it difficult to generate effective advertising audio. Therefore, there is a need for a system that can recognize users' actual emotions in real time and reflect them in audio data.
[0438] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for voice actors to register their own voice data and usage conditions, means for video producers to search for voice actors using keywords, means for converting input text into voice based on the voice data of the selected voice actor, means for recognizing the user's emotions and generating voice data based on those emotions, and means for incorporating the generated voice data into video. This makes it possible to generate emotionally rich voice data that reflects the user's emotions in real time.
[0439] The "voice actor registration means" is a means for a voice actor to register his / her voice data and usage conditions in the system.
[0440] The "voice actor search means" is a means for a video producer to use keywords to search for appropriate voice actors from a database.
[0441] The "voice conversion means" is a means for converting input text into voice based on the voice data of the selected voice actor.
[0442] The "emotion recognition means" is a means for recognizing the user's emotion and generating appropriate voice data based on that emotion.
[0443] The "voice generation means" is a means for converting input text into an emotionally rich voice based on the emotion recognized by the emotion recognition means.
[0444] The "video incorporation means" is a means for integrating the generated audio data into video content to generate a final video work.
[0445] The "database storage means" is a means for storing voice data registered by voice actors and the conditions for using the data in a database.
[0446] This invention relates to a system that recognizes a user's emotions in real time, generates voice data based on the emotions, and incorporates the voice data into video works and advertisements. This system includes a voice actor registration means, a voice actor search means, a voice conversion means, an emotion recognition means, a voice generation means, a video embedding means, and a database storage means.
[0447] composition
[0448] The system mainly consists of the following components:
[0449] 1. Hardware:
[0450] Smartphone or head-mounted display (HMD): Equipped with a camera and microphone to recognize the user's emotions.
[0451] Cloud database: A database (e.g., Google Firebase) for storing voice actor audio data and terms of use.
[0452] 2. Software:
[0453] Emotion recognition engine (Emotion AI): Recognizes emotions from the user's facial expressions and voice in real time.
[0454] Speech generation AI (Text-to-Speech engine): Converts input text into emotive speech.
[0455] Front-end framework (React Native): A mobile application to provide the user interface.
[0456] Program processing
[0457] The server first allows users to access the system and create an account. Voice actors register their voice data and terms of use, which are then stored in a database. This process is carried out through a front-end interface using React Native and is stored in Firebase.
[0458] The filmmaker (user) logs into the system using a smartphone or HMD and inputs specific keywords to search for suitable voice actors. Results are returned from the cloud database, and the user can select from a list of voice actors.
[0459] Based on the voice data of the selected voice actor, the user inputs a prompt, such as advertising text like "Buy now and get a special discount!". At this time, the emotion recognition engine (Emotion AI) analyzes the user's facial expressions and voice and collects emotional data in real time.
[0460] The emotion data collected by the emotion recognition engine is sent to the server, and then the voice generation AI converts the prompt sentence into voice based on this emotion data. For example, if the user expresses the emotion "excitement," voice data of the voice actor that reflects this emotion will be generated.
[0461] The generated audio data is integrated into the user's video content or advertisements using a video integration means, allowing the user to easily create high-quality video works that reflect emotional audio.
[0462] Specific examples
[0463] For example, if an advertiser uses a smartphone app to input the text "Buy now and get a special discount!", Emotion AI will recognize the emotion "excited" from the user's facial expressions and voice. The speech generation AI will then convert the text into speech, which will then be incorporated into the ad.
[0464] Prompt Sentence Examples
[0465] "Buy now and get a special discount!"
[0466] Recognized emotion: "Excitement"
[0467] Use voice data from a voice actor to convert this text into emotive speech.
[0468] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0469] Step 1:
[0470] A user accesses the system and creates an account. The user (voice actor) enters their voice data and terms of use and uploads it to the server. The server validates this information and stores it in a cloud database (e.g., Google Firebase). This creates a voice database that can be searched later.
[0471] Step 2:
[0472] The user (filmmaker) logs into the system and searches for a voice actor by entering a keyword. The entered keyword is sent to the server, which searches the database for a list of matching voice actors. The server displays the search results to the user, who then selects a voice actor from the displayed list.
[0473] Step 3:
[0474] The user (video creator) accesses the details page of the voice actor they selected and enters a prompt sentence (e.g., "Buy now and get a special discount!"). This entered text is sent to the server. At the same time, the emotion recognition engine (Emotion AI) collects emotional data from the user's facial expressions and voice in real time and sends it to the server.
[0475] Step 4:
[0476] The server sends prompts to a text-to-speech engine based on the acquired emotional data and the input text. The text-to-speech engine then generates appropriate, emotionally rich speech data. In this process, the emotional data sent from the emotion recognition engine is reflected in the speech generation.
[0477] Step 5:
[0478] The server converts the generated audio data into an appropriate format and provides it to the user's (video creator's) device. The user then downloads the generated audio data and integrates it into their own video content or advertisements. This results in the creation of a high-quality video work that includes richly expressive audio.
[0479] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0480] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0481] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0482] [Second embodiment]
[0483] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0484] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0485] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0486] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0487] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0488] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0489] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0490] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0491] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0492] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0493] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0494] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0495] The present invention relates to a system for easily generating audio data and incorporating it into a self-produced animation or test video. Specific embodiments of the system of the present invention will be described below.
[0496] System Overview
[0497] This system consists of three main components.
[0498] 1. A means for voice actors to register their own voice data and terms of use
[0499] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[0500] 3. Means for converting input text into speech based on the voice data of the selected voice actor and incorporating the generated speech data into the video.
[0501] Program processing
[0502] 1. Voice actor voice data registration
[0503] A user (voice actor) accesses the website and creates an account.
[0504] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[0505] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[0506] The server validates the uploaded audio data and stores it in a database.
[0507] 2. Search and select a voice actor
[0508] The user (video creator) logs in to the website and accesses the voice actor search page.
[0509] The user (video creator) enters keywords and performs a search.
[0510] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[0511] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0512] 3. Generating audio data
[0513] Enter text on the detailed information page of the voice actor selected by the user (video producer).
[0514] The server uses speech generation AI to convert the input text into audio data.
[0515] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[0516] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[0517] Specific examples
[0518] For example, consider the case where a user (filmmaker) is looking for a "young female voice" to use in one of their own animation works. The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the results. The user selects an appropriate voice actor from the results and enters the text "Hello, nice to meet you" on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed easily.
[0519] In this way, the system of the present invention provides an environment in which users can quickly and efficiently generate audio data and use it in video productions, significantly reducing production costs and time and enabling even amateur producers to achieve professional-looking results.
[0520] The processing flow will be explained below.
[0521] Voice actor voice data registration process
[0522] Step 1:
[0523] The user (voice actor) accesses the website and opens the account creation page.
[0524] The device presents the user with an account creation form.
[0525] Step 2:
[0526] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[0527] The terminal sends the input information to the server.
[0528] Step 3:
[0529] The server creates a new user record in the database based on the received account information.
[0530] The server validates the input and saves it to the database.
[0531] Step 4:
[0532] The server sends a notification to the user that account registration is complete.
[0533] The server sends emails and site notifications.
[0534] Step 5:
[0535] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[0536] The device displays a voice registration form.
[0537] Step 6:
[0538] The user (voice actor) fills out the required information in the form and uploads the audio file.
[0539] The device sends voice data and other input information to the server.
[0540] Step 7:
[0541] The server validates the uploaded audio data and stores it in the database.
[0542] The server checks the format and size of the audio data and stores it in a database.
[0543] Step 8:
[0544] The server sends a notification to the user that the voice data has been registered.
[0545] The server sends emails and site notifications.
[0546] Voice Actor Search and Selection Process
[0547] Step 1:
[0548] The user (video creator) logs in to the website and accesses the voice actor search page.
[0549] The terminal displays a login form and, after successful authentication, a search page.
[0550] Step 2:
[0551] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[0552] The terminal sends the entered keyword to the server.
[0553] Step 3:
[0554] The server searches the database for voice actors that match the keywords.
[0555] The server executes a search query against the database and retrieves the results.
[0556] Step 4:
[0557] The server sends the search results to the user's (filmmaker's) device.
[0558] The server sends the search results in JSON format, which the device parses for display.
[0559] Step 5:
[0560] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0561] The device will display a detailed information page.
[0562] Voice data generation process
[0563] Step 1:
[0564] On the detailed information page of the voice actor selected by the user (video producer), enter a line (e.g., "Hello, nice to meet you") in the text input box and click the "Generate Voice" button.
[0565] The device sends the entered text and the selected voice actor's ID to the server.
[0566] Step 2:
[0567] The server sends a voice generation request to the voice generation AI.
[0568] The server sends the text and voice actor ID to the voice generation API and receives the voice data.
[0569] Step 3:
[0570] The server converts the generated audio data into a specific format and sends it back to the user (video producer).
[0571] The server converts the audio data into MP3 or WAV format and sends it to the device.
[0572] Step 4:
[0573] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[0574] The device displays a download link for the audio data, which the user can then download and use.
[0575] Specific examples
[0576] For example, a user (video creator) searches for "the voice of a young woman," selects a specific voice actor, enters the text "Hello, nice to meet you," and generates the voice. Through this process, a series of steps are realized in which the generated voice data is received from the server, downloaded, and incorporated into the user's own animated work.
[0577] Example 1
[0578] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0579] In today's video production, generating and incorporating audio data requires a significant amount of time and cost. Searching, selecting, and managing audio from professional voice actors is particularly complex, and the lack of an efficient system is a problem. This makes amateur productions and short-term projects difficult.
[0580] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0581] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for incorporating the generated voice data into a video work, a means for converting input text into voice using a voice generation AI model, and a means for providing the generated voice data to users, thereby enabling video producers to efficiently generate and incorporate voice data.
[0582] A "voice actor" is an individual or group that uses their own voice to provide audio data.
[0583] "Audio data" refers to audio information provided by voice actors stored in digital data format.
[0584] "Conditions of use" refers to restrictions and permissions regarding the use of audio data, such as whether commercial use is permitted or not, and the expiration date of use.
[0585] "Filmmaker" means an individual or organization that produces a film work.
[0586] A "keyword" is a character string or phrase that a user enters to search for a voice actor.
[0587] A "voice generation AI model" is a type of artificial intelligence technology that analyzes input text and generates it as voice data.
[0588] "Text" refers to the written information that is the basis for generating speech.
[0589] A "server" is a computer system that processes various types of data and provides services to users.
[0590] A "database" is a system that systematically stores and manages information such as voice data and usage conditions.
[0591] "Prompt sentence" refers to an example of text that a user inputs into a speech-generation AI model.
[0592] This invention relates to a system that allows filmmakers to efficiently generate audio data and incorporate it into their video works. This system has the means to register, search, generate, and provide audio data, and by using a generative AI model, it realizes rapid and high-quality audio data generation.
[0593] System Program Overview
[0594] The program of this system is implemented using the following hardware and software.
[0595] Server: A computer system for storing, searching, generating, and providing data.
[0596] Database: Management of voice data, terms of use, user information, etc.
[0597] Website: The interface for users (voice actors and filmmakers) to access
[0598] Speech generation AI model: Artificial intelligence technology that analyzes input text and generates voice data
[0599] Voice data registration details
[0600] A user (voice actor) accesses the website and creates an account. When creating an account, the user enters the required information and clicks "Register," and the server saves the information in a database. Once registration is complete, the user (voice actor) logs in and registers their voice data, terms of use, and fees. Once the voice data is uploaded, the server validates it and saves it in a database.
[0601] More on finding and selecting voice actors
[0602] The user (filmmaker) logs in to the website and accesses the voice actor search page. Using the search function, they enter a keyword, such as "female female voice," and execute a search. The server searches the database for the relevant voice actor and presents the results to the user (filmmaker). The user can then check the detailed information from the displayed list of voice actors and select the one they need.
[0603] Voice data generation details
[0604] The user (video producer) enters text, such as lines, on the detailed information page of the voice actor they selected. When they click the "Generate" button, the server uses a voice generation AI model to convert the entered text into audio data. The generated audio data is then converted by the server into an appropriate format and provided to the user. The user (video producer) can then download this audio data and incorporate it into their own video work.
[0605] Specific examples
[0606] For example, if a user (filmmaker) is looking for a "young female voice" to use in their own animation work, they can proceed as follows: The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the search results. The user selects an appropriate voice actor from the list and enters text, such as "Hello, nice to meet you," on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed.
[0607] Example prompt sentence:
[0608] "Please convert the text "Hello, nice to meet you" into speech data in a young female voice."
[0609] In this way, the system of the present invention is a mechanism that realizes efficient and high-quality audio data generation and provides video producers with audio data that can be used quickly.
[0610] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0611] Step 1:
[0612] User (voice actor) account creation and information registration
[0613] The user (voice actor) accesses the website and enters the necessary information into the "Create Account" form, such as name, email address, and password.
[0614] When a user clicks "Register," the server receives the information and performs validation, such as checking the format of the email address and the strength of the password.
[0615] If validation passes, the server saves the account information in the database, allowing the user to log in.
[0616] After logging in, the user (voice actor) registers their own voice data, terms of use, and usage fees. After uploading the voice data, the server validates the data and stores it in the database. Specifically, it checks the format and size of the voice file.
[0617] Input: Account information entered by the user (name, email address, password) and voice data
[0618] Output: Account information and voice data stored in a database
[0619] Step 2:
[0620] User (video creator) login and voice actor search
[0621] The user (video creator) logs in to the website and accesses the voice actor search page. A username and password are required to log in.
[0622] The user (filmmaker) enters a keyword into the search bar, for example, "young woman's voice," and executes the search.
[0623] The server searches the database for a list of matching voice actors and retrieves the results, including the voice data and usage conditions that match the search criteria.
[0624] The server displays the search results to the user (video creator), which include the name of the voice actor, sample audio, terms of use, etc.
[0625] Input: The keyword that the user types into the search bar
[0626] Output: A list of voice actors displayed as search results
[0627] Step 3:
[0628] User (video creator) voice actor selection and text input
[0629] The user (filmmaker) selects the appropriate voice actor from the search results and accesses their details page.
[0630] The user checks the details and enters the text they want to convert into speech in the text entry field, for example, "Hello, nice to meet you."
[0631] The user clicks the "Generate" button, which sends the entered text to the server.
[0632] Input: The voice actor the user selects, and the text they enter into the text field.
[0633] Output: Text data sent to the server
[0634] Step 4:
[0635] Voice data generation by the server
[0636] The server sends the input text to a speech generation AI model, which analyzes the text and generates appropriate speech data.
[0637] The voice generation AI model generates speech based on the characteristics of a specified voice actor, for example, generating "Hello, nice to meet you" in a young female voice.
[0638] The server receives the generated audio data and performs format conversion and quality checks, including encoding the audio data and removing noise.
[0639] Input: Text data sent to the server
[0640] Output: Generated audio data
[0641] Step 5:
[0642] Providing audio data to users (video producers)
[0643] The server generates a download link for the generated audio data and provides the link to the user (video producer).
[0644] The user (video creator) downloads the audio data from the provided link and can incorporate the downloaded audio data into their own video work.
[0645] Input: Generated audio data
[0646] Output: Download link and audio data provided to the user
[0647] (Application example 1)
[0648] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0649] In existing content distribution systems, generating audio data and incorporating it in real time is extremely time-consuming, especially when providing an interactive experience. Furthermore, there is a lack of efficient means for audiovisual producers to search for and use the specific audio data they require. This creates a need for technology that can quickly and efficiently generate high-quality audio data and provide interactive experiences through visual display devices in real time.
[0650] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0651] In this invention, the server includes means for audio providers to register their own audio data and usage conditions, means for content creators to search for audio providers using keywords, means for converting input text into audio based on the audio data of the selected audio provider, means for incorporating the generated audio data into content in real time, and means for providing an interactive experience through a visual display device, thereby making it possible to quickly and efficiently generate audio data and provide an interactive experience through a visual display device in real time.
[0652] "Audio Data" means a digital file of sound recorded or generated by an audio provider.
[0653] "Audio provider" is a person or organization that provides audio data and sets the terms of use.
[0654] "Conditions of Use" are specific conditions and restrictions set by the audio provider when using audio data.
[0655] A "content creator" is a person or organization that creates content such as video or audio.
[0656] "Keywords" are specific words or phrases used in searches.
[0657] A "generative AI model" is an artificial intelligence technology that takes input text and converts it into voice data.
[0658] "Real time" refers to a state in which processing and responses occur almost simultaneously.
[0659] A "visual display device" is a device that allows a user to visually view content, and includes smartphones and head-mounted displays.
[0660] An "interactive experience" is one in which a user can interact with a system and receive responses in real time.
[0661] "Content" refers to digital media that includes information such as video and audio.
[0662] The present invention provides a system for generating audio data and incorporating it into content in real time, and is specifically implemented with the following configuration.
[0663] 1. System Overview
[0664] The system includes a voice provider, a content creator, a server, and a visual display device. The voice provider registers their voice data and terms of use, and the content creator searches for the voice provider using keywords, inputs text, and uses a voice generation AI model to generate voice data, which is then incorporated into the content in real time.
[0665] 2. Hardware and Software
[0666] Hardware: Smartphones, head-mounted displays, servers
[0667] Software: Web-based applications, databases (e.g., MySQL), speech generation AI models (e.g., Tacotron, WaveNet)
[0668] 3. System Processing Procedures
[0669] 1. Voice provider voice data registration
[0670] The voice provider visits the website and creates an account.
[0671] Once you enter your account information and register, the server stores that information in a database.
[0672] After completing the registration, the voice provider inputs his / her voice data, the terms of use, and the usage fee, and transmits them to the server.
[0673] The server validates the uploaded audio data and stores it in a database.
[0674] 2. Search and select an audio provider
[0675] A content creator logs into a website and visits the search page.
[0676] A content creator enters keywords and performs a search.
[0677] The server searches the database for the appropriate audio provider and presents the results to the content creator.
[0678] The content creator checks the detailed information from the displayed list of audio providers and selects one.
[0679] 3. Generating audio data
[0680] Enter text on the details page of the audio provider that the content creator has selected.
[0681] The server uses a generative AI model to convert the input text into audio data.
[0682] After successful speech generation, the server converts the generated speech data into an appropriate format and provides an interactive experience through a visual display device in real time.
[0683] 4. Specific Examples
[0684] For example, if a content creator searches for "young female voice" and enters the text "Hello, nice to meet you," the server will use the generative AI model to convert this text into speech data and incorporate the generated speech data into the content in real time, thereby providing an interactive experience to the user through a visual display device.
[0685] 5. Examples of prompts
[0686] Generative AI models:
[0687] Prompt: Convert the following text into the voice of the voice actor you provided: "Hello, nice to meet you."
[0688] Expected output: The text "Hello, nice to meet you" will be output as audio data from the selected young female voice actress.
[0689] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0690] Step 1:
[0691] The voice provider accesses the website and creates an account. The account information entered by the voice provider is sent to the server, which stores the information in a database.
[0692] Input: Account information input by the voice provider
[0693] Data processing: The server receives the account information and stores it in the database in the appropriate format.
[0694] Output: Account registration completion notification
[0695] Step 2:
[0696] After completing registration, the voice provider enters their voice data, terms of use, and usage fees into a form on the website and submits it to the server, which then validates the uploaded voice data and stores it in a database.
[0697] Input: Input of audio data, terms of use, and fees by audio provider
[0698] Data processing: The server receives the data, validates the voice data, and saves it to the database.
[0699] Output: Notification of completion of voice data registration
[0700] Step 3:
[0701] A content creator logs in to the website and accesses the search page. The content creator enters keywords and performs a search. The server searches the database for the corresponding audio provider and presents the results to the content creator.
[0702] Input: Keywords entered by content creators
[0703] Data processing: The server performs a keyword search and extracts the corresponding voice providers from the database.
[0704] Output: Presenting search results
[0705] Step 4:
[0706] The content creator checks the details of the audio providers from the list displayed and selects one. On the details page of the selected audio provider, they input text and request audio generation.
[0707] Input: Content creator selects audio provider and inputs text
[0708] Data processing: The server receives the text and invokes a generative AI model based on the voice data of the selected voice provider.
[0709] Output: Notification of execution of generation request
[0710] Step 5:
[0711] The server uses the generative AI model to convert the input text into speech data, and after successful speech generation, converts the generated speech data into an appropriate format and incorporates it into the content in real time.
[0712] Input: Text input to the generative AI model
[0713] Data processing: The generative AI model converts the text into audio data, and the server retrieves the audio data.
[0714] Output: Generated audio data
[0715] Step 6:
[0716] The generated audio data is incorporated into the content through a visual display device to provide an interactive experience for the user.
[0717] Input: Generated audio data
[0718] Data processing: Integrating audio data into the content of a visual display
[0719] Output: Interactive experience through a visual display device
[0720] This series of processing steps allows content creators to efficiently generate audio data and deliver real-time interactive experiences through visual display devices.
[0721] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0722] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user-created animation or test video. A specific embodiment of the system of the present invention will be described below.
[0723] System Overview
[0724] This system consists of four main components.
[0725] 1. A means for voice actors to register their own voice data and terms of use
[0726] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[0727] 3. Means for converting input text into speech based on the voice data of the selected voice actor.
[0728] 4. A means to recognize the user's emotions using an emotion engine and generate voice data based on that information
[0729] Program processing
[0730] 1. Voice actor voice data registration
[0731] A user (voice actor) accesses the website and creates an account.
[0732] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[0733] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[0734] The server validates the uploaded audio data and stores it in a database.
[0735] 2. Search and select a voice actor
[0736] The user (video creator) logs in to the website and accesses the voice actor search page.
[0737] The user (video creator) enters keywords and performs a search.
[0738] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[0739] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0740] 3. Emotion-Recognition Assistance
[0741] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[0742] The device recognizes the user's emotions in real time from their facial expressions and voice, and sends that information to the server.
[0743] The emotion engine assists in generating voice data containing appropriate intonation and emotional expression based on the recognized emotion.
[0744] 4. Generating Audio Data
[0745] The server uses a speech generation AI and emotion engine to convert the input text into voice data.
[0746] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[0747] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[0748] Specific examples
[0749] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. In this process, the emotion "joy" is recognized from the user's facial expressions and voice, and the server selects voice data from the voice actor based on that, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[0750] In this way, the system of the present invention can significantly improve the quality of video works by recognizing the user's emotions in real time and providing appropriate audio data, thereby significantly reducing production costs and time and providing an easy-to-use environment for users to easily create high-quality works.
[0751] The processing flow will be explained below.
[0752] Voice actor voice data registration process
[0753] Step 1:
[0754] The user (voice actor) accesses the website and opens the account creation page.
[0755] The device presents the user with an account creation form.
[0756] Step 2:
[0757] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[0758] The terminal sends the input information to the server.
[0759] Step 3:
[0760] The server creates a new user record in the database based on the received account information.
[0761] The server validates the input and saves it to the database.
[0762] Step 4:
[0763] The server sends a notification to the user that account registration is complete.
[0764] The server sends emails and site notifications.
[0765] Step 5:
[0766] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[0767] The device displays a voice registration form.
[0768] Step 6:
[0769] The user (voice actor) fills out the required information in the form and uploads the audio file.
[0770] The device sends voice data and other input information to the server.
[0771] Step 7:
[0772] The server validates the uploaded audio data and stores it in the database.
[0773] The server checks the format and size of the audio data and stores it in a database.
[0774] Step 8:
[0775] The server sends a notification to the user that the voice data has been registered.
[0776] The server sends emails and site notifications.
[0777] Voice Actor Search and Selection Process
[0778] Step 1:
[0779] The user (video creator) logs in to the website and accesses the voice actor search page.
[0780] The terminal displays a login form and, after successful authentication, a search page.
[0781] Step 2:
[0782] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[0783] The terminal sends the entered keyword to the server.
[0784] Step 3:
[0785] The server searches the database for voice actors that match the keywords.
[0786] The server executes a search query against the database and retrieves the results.
[0787] Step 4:
[0788] The server sends the search results to the user's (filmmaker's) device.
[0789] The server sends the search results in JSON format, which the device parses for display.
[0790] Step 5:
[0791] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0792] The device will display a detailed information page.
[0793] Assisted processes using emotion recognition
[0794] Step 1:
[0795] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[0796] The terminal displays an input text box.
[0797] Step 2:
[0798] The user (video creator) uses a camera and microphone to provide facial expressions and voice to the system.
[0799] The device captures the user's facial expressions and voice data in real time and sends it to the server.
[0800] Step 3:
[0801] The server uses an emotion engine to recognize the user's emotions in real time from the transmitted facial and voice data.
[0802] The server sends the input data to the emotion engine and analyzes the emotions.
[0803] Step 4:
[0804] Based on the recognized emotion, the server sends a request to the speech generation AI to generate speech data containing expressions that best fit the input text.
[0805] The server sends the emotion information and text to the speech generation AI.
[0806] Step 5:
[0807] The server receives the generated audio data, converts it into an appropriate format, and provides it to the user (video producer).
[0808] The server converts the audio data into MP3 or WAV format and sends it to the device.
[0809] Step 6:
[0810] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[0811] The device displays a download link for the audio data, which the user can then download and use.
[0812] Specific examples
[0813] For example, a user (video producer) searches for "a young female voice," selects a specific voice actor, and enters the text "Hello, nice to meet you." Next, the user uses a camera or microphone to provide the system with their own smiling face and cheerful voice, which the emotion engine recognizes as "joy." Based on this emotional information, the server sends a request to the voice generation AI, which generates voice data reflecting the emotion of "joy." The user can download this generated voice data and incorporate it into their own animated work, enabling them to create an emotionally rich production.
[0814] Example 2
[0815] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0816] Conventionally, when generating audio data in video production, it has been difficult to provide audio that reflects the user's emotions. Furthermore, searching and managing voice data for voice actors is cumbersome, increasing production costs and time. This has led to the problem of making it difficult to efficiently produce high-quality video works.
[0817] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0818] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for recognizing the user's emotions and generating voice data based on that information, and a means for incorporating the generated voice data into video. This makes it possible to efficiently generate high-quality voice data that reflects the user's emotions and incorporate it into video.
[0819] "Voice actors" are people who register their own voice data and play the role of providing audio for video production.
[0820] "Video producers" are people who generate the necessary audio data and perform editing work when producing a video work.
[0821] "Audio Data" means audio recordings provided by voice actors or audio generated by a generative AI model.
[0822] "Terms of Use" refers to the conditions and restrictions for using the voice data provided by the voice actors.
[0823] "Keywords" are words or phrases that filmmakers use to search for voice actors.
[0824] "Emotion recognition" is a technology that analyzes a user's facial expressions and voice and identifies their emotions in real time.
[0825] "Speech generation artificial intelligence" is a technology that generates natural-sounding speech based on input text and emotional data.
[0826] The "database" is an information collection system for efficiently managing and storing voice actor audio data and usage conditions.
[0827] "Text" refers to the string of characters entered by the video producer, and is the words or sentences that form the basis of the audio data.
[0828] "Video embedding" refers to the process of embedding the generated audio data as part of a video production.
[0829] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user's own animation or video work. A specific embodiment of the system of the present invention is described below.
[0830] This system consists of four main components.
[0831] 1. How to register voice data
[0832] A user (voice actor) accesses the website and creates an account by entering information such as name, email address, and password, and submitting it.
[0833] The server validates the entered information and saves it in the database. After successful saving, the user (voice actor) enters the voice data, terms of use, and fees, and uploads it.
[0834] The server validates the audio data and stores it in the database.
[0835] 2. How to search for voice actors
[0836] The user (video creator) logs in to the system and accesses the voice actor search page, where they enter keywords and perform a search.
[0837] The server searches the database for the relevant voice actor and displays the results as a list to the user (video producer).
[0838] The user (video creator) clicks and selects detailed information from the list of voice actors displayed.
[0839] 3. Text Input and Emotion Recognition
[0840] The user (video producer) goes to the details page of the voice actor they have chosen and enters the lines in the text input box.
[0841] The user's facial expressions and voice are captured in real time, and the device recognizes their emotions using emotion recognition software (e.g., OpenFace or Microsoft Emotion API).
[0842] The recognized emotion data is sent to a server.
[0843] 4. Audio data generation means
[0844] The server uses a speech generation AI (for example, Google Text-to-Speech API or Amazon Polly) to generate speech based on the input text and emotion data.
[0845] The server converts the generated audio data into an appropriate format and provides it to the user (video producer).
[0846] The user (video creator) downloads the generated audio data and incorporates it into their own video work.
[0847] Specific examples
[0848] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. During this process, the emotion "joy" is recognized from the user's facial expressions and voice. Based on this, the server selects voice data from the voice actor, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[0849] An example of a prompt to input to a generative AI model is as follows:
[0850] 1. "In a young female voice, please vocalize the lines "Hello" and "Nice to meet you" with a joyful emotion."
[0851] 2. "Can you say 'I'm looking forward to tomorrow!' in an energetic male voice with a surprised expression?"
[0852] 3. "Generate a recording of a middle-aged man quietly saying "goodbye" with sadness in his voice."
[0853] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[0854] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0855] Step 1: Create an account (a method for registering audio data)
[0856] Input: The user (voice actor) visits the website and enters their account information, including their name, email address, and password.
[0857] Data Processing: The server validates the entered information to ensure accuracy and format.
[0858] Output: If validation is successful, the server saves the account information to the database and sends the user a confirmation email.
[0859] Specific operation: The user (voice actor) enters the required information into the web form and clicks the "Create Account" button. The server saves the information in the database and automatically sends a confirmation email.
[0860] Step 2: Upload the audio data and terms of use
[0861] Input: After logging in, the user (voice actor) enters the audio data, terms of use, and usage fee into the web form and presses the upload button.
[0862] Data processing: The server receives the audio file and validates the file format and quality. It also checks the terms of use and fees.
[0863] Output: The voice data and terms of use that pass validation are saved in the database, and a success notification is displayed to the user.
[0864] Specific operation: The user (voice actor) selects an audio file, enters the required conditions, and clicks the "Upload" button. The server processes this and notifies the user of the result.
[0865] Step 3: Find a voice actor
[0866] Input: The user (video producer) logs into the system, accesses the voice actor search page, enters keywords, and presses the search button.
[0867] Data calculation: The server searches the database for the appropriate voice actor and filters the results that match the keywords.
[0868] Output: The search results are displayed as a list, and detailed information about the corresponding voice actors is presented to the user.
[0869] Specific operation: The user (filmmaker) enters keywords in the search box and clicks the "Search" button. The server retrieves the results and displays them on the page.
[0870] Step 4: Select a voice actor and enter text
[0871] Input: The user (video creator) clicks on the displayed voice actor details page and enters lines in the text input box.
[0872] Data processing: The server records the user's selection and temporarily stores the text data.
[0873] Output: A web page displays text entry boxes and other details that the user types and are sent to the server and stored.
[0874] Specific operation: The user (filmmaker) opens the details page, enters the dialogue text, and clicks the "Submit" button. The server receives and stores this information.
[0875] Step 5: Emotion Recognition
[0876] Input: While the user (video creator) is entering text, the device captures the user's facial expressions and voice in real time.
[0877] Data calculation: The device uses emotion recognition software (e.g., OpenFace or Microsoft Emotion API) to analyze the captured data and generate emotion data.
[0878] Output: The generated emotion data is sent to the server and associated with the text data.
[0879] How it works: While the user (filmmaker) is typing text, the device captures facial expressions and voice using the built-in camera and microphone, and runs emotion recognition software.
[0880] Step 6: Generate audio data
[0881] Input: The server sends the input text and emotion data to a speech generation AI (e.g., Google Text-to-Speech API or Amazon Polly).
[0882] Data calculation: The voice generation AI processes the input data based on the prompt sentence and generates voice data that reflects the emotion.
[0883] Output: The generated audio data is returned to the server and converted into the appropriate format.
[0884] Specific operation: The server passes text and emotion data to the speech generation AI, receives the generated speech data, and performs format conversion.
[0885] Step 7: Provide audio data
[0886] Input: The server provides the generated audio data to the user (video producer).
[0887] Data processing: The server rechecks the quality of the audio data and generates a download link.
[0888] Output: A page is displayed containing a link that allows the user (the filmmaker) to download the audio data.
[0889] Specific operation: The user (video creator) clicks on the download link, obtains the audio data, and incorporates it into the video work.
[0890] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[0891] (Application example 2)
[0892] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0893] A major challenge in current video production is the significant time and expense required to generate appropriate audio data. Particularly in the advertising field, emotive audio resonates with viewers and increases the effectiveness of advertising. However, existing speech synthesis technologies have limitations in expressing emotions, making it difficult to generate effective advertising audio. Therefore, there is a need for a system that can recognize users' actual emotions in real time and reflect them in audio data.
[0894] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for voice actors to register their own voice data and usage conditions, means for video producers to search for voice actors using keywords, means for converting input text into voice based on the voice data of the selected voice actor, means for recognizing the user's emotions and generating voice data based on those emotions, and means for incorporating the generated voice data into video. This makes it possible to generate emotionally rich voice data that reflects the user's emotions in real time.
[0895] The "voice actor registration means" is a means for a voice actor to register his / her voice data and usage conditions in the system.
[0896] The "voice actor search means" is a means for a video producer to use keywords to search for appropriate voice actors from a database.
[0897] The "voice conversion means" is a means for converting input text into voice based on the voice data of the selected voice actor.
[0898] The "emotion recognition means" is a means for recognizing the user's emotion and generating appropriate voice data based on that emotion.
[0899] The "voice generation means" is a means for converting input text into an emotionally rich voice based on the emotion recognized by the emotion recognition means.
[0900] The "video incorporation means" is a means for integrating the generated audio data into video content to generate a final video work.
[0901] The "database storage means" is a means for storing voice data registered by voice actors and the conditions for using the data in a database.
[0902] This invention relates to a system that recognizes a user's emotions in real time, generates voice data based on the emotions, and incorporates the voice data into video works and advertisements. This system includes a voice actor registration means, a voice actor search means, a voice conversion means, an emotion recognition means, a voice generation means, a video embedding means, and a database storage means.
[0903] composition
[0904] The system mainly consists of the following components:
[0905] 1. Hardware:
[0906] Smartphone or head-mounted display (HMD): Equipped with a camera and microphone to recognize the user's emotions.
[0907] Cloud database: A database (e.g., Google Firebase) for storing voice actor audio data and terms of use.
[0908] 2. Software:
[0909] Emotion recognition engine (Emotion AI): Recognizes emotions from the user's facial expressions and voice in real time.
[0910] Speech generation AI (Text-to-Speech engine): Converts input text into emotive speech.
[0911] Front-end framework (React Native): A mobile application to provide the user interface.
[0912] Program processing
[0913] The server first allows users to access the system and create an account. Voice actors register their voice data and terms of use, which are then stored in a database. This process is carried out through a front-end interface using React Native and is stored in Firebase.
[0914] The filmmaker (user) logs into the system using a smartphone or HMD and inputs specific keywords to search for suitable voice actors. Results are returned from the cloud database, and the user can select from a list of voice actors.
[0915] Based on the voice data of the selected voice actor, the user inputs a prompt, such as advertising text like "Buy now and get a special discount!". At this time, the emotion recognition engine (Emotion AI) analyzes the user's facial expressions and voice and collects emotional data in real time.
[0916] The emotion data collected by the emotion recognition engine is sent to the server, and then the voice generation AI converts the prompt sentence into voice based on this emotion data. For example, if the user expresses the emotion "excitement," voice data of the voice actor that reflects this emotion will be generated.
[0917] The generated audio data is integrated into the user's video content or advertisements using a video integration means, allowing the user to easily create high-quality video works that reflect emotional audio.
[0918] Specific examples
[0919] For example, if an advertiser uses a smartphone app to input the text "Buy now and get a special discount!", Emotion AI will recognize the emotion "excited" from the user's facial expressions and voice. The speech generation AI will then convert the text into speech, which will then be incorporated into the ad.
[0920] Prompt Sentence Examples
[0921] "Buy now and get a special discount!"
[0922] Recognized emotion: "Excitement"
[0923] Use voice data from a voice actor to convert this text into emotive speech.
[0924] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0925] Step 1:
[0926] A user accesses the system and creates an account. The user (voice actor) enters their voice data and terms of use and uploads it to the server. The server validates this information and stores it in a cloud database (e.g., Google Firebase). This creates a voice database that can be searched later.
[0927] Step 2:
[0928] The user (filmmaker) logs into the system and searches for a voice actor by entering a keyword. The entered keyword is sent to the server, which searches the database for a list of matching voice actors. The server displays the search results to the user, who then selects a voice actor from the displayed list.
[0929] Step 3:
[0930] The user (video creator) accesses the details page of the voice actor they selected and enters a prompt sentence (e.g., "Buy now and get a special discount!"). This entered text is sent to the server. At the same time, the emotion recognition engine (Emotion AI) collects emotional data from the user's facial expressions and voice in real time and sends it to the server.
[0931] Step 4:
[0932] The server sends prompts to a text-to-speech engine based on the acquired emotional data and the input text. The text-to-speech engine then generates appropriate, emotionally rich speech data. In this process, the emotional data sent from the emotion recognition engine is reflected in the speech generation.
[0933] Step 5:
[0934] The server converts the generated audio data into an appropriate format and provides it to the user's (video creator's) device. The user then downloads the generated audio data and integrates it into their own video content or advertisements. This results in the creation of a high-quality video work that includes richly expressive audio.
[0935] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0936] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0937] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0938] [Third embodiment]
[0939] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0940] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0941] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0942] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0943] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0944] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0945] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0946] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0947] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0948] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0949] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0950] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0951] The present invention relates to a system for easily generating audio data and incorporating it into a self-produced animation or test video. Specific embodiments of the system of the present invention will be described below.
[0952] System Overview
[0953] This system consists of three main components.
[0954] 1. A means for voice actors to register their own voice data and terms of use
[0955] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[0956] 3. Means for converting input text into speech based on the voice data of the selected voice actor and incorporating the generated speech data into the video.
[0957] Program processing
[0958] 1. Voice actor voice data registration
[0959] A user (voice actor) accesses the website and creates an account.
[0960] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[0961] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[0962] The server validates the uploaded audio data and stores it in a database.
[0963] 2. Search and select a voice actor
[0964] The user (video creator) logs in to the website and accesses the voice actor search page.
[0965] The user (video creator) enters keywords and performs a search.
[0966] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[0967] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[0968] 3. Generating audio data
[0969] Enter text on the detailed information page of the voice actor selected by the user (video producer).
[0970] The server uses speech generation AI to convert the input text into audio data.
[0971] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[0972] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[0973] Specific examples
[0974] For example, consider the case where a user (filmmaker) is looking for a "young female voice" to use in one of their own animation works. The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the results. The user selects an appropriate voice actor from the results and enters the text "Hello, nice to meet you" on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed easily.
[0975] In this way, the system of the present invention provides an environment in which users can quickly and efficiently generate audio data and use it in video productions, significantly reducing production costs and time and enabling even amateur producers to achieve professional-looking results.
[0976] The processing flow will be explained below.
[0977] Voice actor voice data registration process
[0978] Step 1:
[0979] The user (voice actor) accesses the website and opens the account creation page.
[0980] The device presents the user with an account creation form.
[0981] Step 2:
[0982] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[0983] The terminal sends the input information to the server.
[0984] Step 3:
[0985] The server creates a new user record in the database based on the received account information.
[0986] The server validates the input and saves it to the database.
[0987] Step 4:
[0988] The server sends a notification to the user that account registration is complete.
[0989] The server sends emails and site notifications.
[0990] Step 5:
[0991] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[0992] The device displays a voice registration form.
[0993] Step 6:
[0994] The user (voice actor) fills out the required information in the form and uploads the audio file.
[0995] The device sends voice data and other input information to the server.
[0996] Step 7:
[0997] The server validates the uploaded audio data and stores it in the database.
[0998] The server checks the format and size of the audio data and stores it in a database.
[0999] Step 8:
[1000] The server sends a notification to the user that the voice data has been registered.
[1001] The server sends emails and site notifications.
[1002] Voice Actor Search and Selection Process
[1003] Step 1:
[1004] The user (video creator) logs in to the website and accesses the voice actor search page.
[1005] The terminal displays a login form and, after successful authentication, a search page.
[1006] Step 2:
[1007] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[1008] The terminal sends the entered keyword to the server.
[1009] Step 3:
[1010] The server searches the database for voice actors that match the keywords.
[1011] The server executes a search query against the database and retrieves the results.
[1012] Step 4:
[1013] The server sends the search results to the user's (filmmaker's) device.
[1014] The server sends the search results in JSON format, which the device parses for display.
[1015] Step 5:
[1016] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1017] The device will display a detailed information page.
[1018] Voice data generation process
[1019] Step 1:
[1020] On the detailed information page of the voice actor selected by the user (video producer), enter a line (e.g., "Hello, nice to meet you") in the text input box and click the "Generate Voice" button.
[1021] The device sends the entered text and the selected voice actor's ID to the server.
[1022] Step 2:
[1023] The server sends a voice generation request to the voice generation AI.
[1024] The server sends the text and voice actor ID to the voice generation API and receives the voice data.
[1025] Step 3:
[1026] The server converts the generated audio data into a specific format and sends it back to the user (video producer).
[1027] The server converts the audio data into MP3 or WAV format and sends it to the device.
[1028] Step 4:
[1029] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[1030] The device displays a download link for the audio data, which the user can then download and use.
[1031] Specific examples
[1032] For example, a user (video creator) searches for "the voice of a young woman," selects a specific voice actor, enters the text "Hello, nice to meet you," and generates the voice. Through this process, a series of steps are realized in which the generated voice data is received from the server, downloaded, and incorporated into the user's own animated work.
[1033] Example 1
[1034] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1035] In today's video production, generating and incorporating audio data requires a significant amount of time and cost. Searching, selecting, and managing audio from professional voice actors is particularly complex, and the lack of an efficient system is a problem. This makes amateur productions and short-term projects difficult.
[1036] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1037] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for incorporating the generated voice data into a video work, a means for converting input text into voice using a voice generation AI model, and a means for providing the generated voice data to users, thereby enabling video producers to efficiently generate and incorporate voice data.
[1038] A "voice actor" is an individual or group that uses their own voice to provide audio data.
[1039] "Audio data" refers to audio information provided by voice actors stored in digital data format.
[1040] "Conditions of use" refers to restrictions and permissions regarding the use of audio data, such as whether commercial use is permitted or not, and the expiration date of use.
[1041] "Filmmaker" means an individual or organization that produces a film work.
[1042] A "keyword" is a character string or phrase that a user enters to search for a voice actor.
[1043] A "voice generation AI model" is a type of artificial intelligence technology that analyzes input text and generates it as voice data.
[1044] "Text" refers to the written information that is the basis for generating speech.
[1045] A "server" is a computer system that processes various types of data and provides services to users.
[1046] A "database" is a system that systematically stores and manages information such as voice data and usage conditions.
[1047] "Prompt sentence" refers to an example of text that a user inputs into a speech-generation AI model.
[1048] This invention relates to a system that allows filmmakers to efficiently generate audio data and incorporate it into their video works. This system has the means to register, search, generate, and provide audio data, and by using a generative AI model, it realizes rapid and high-quality audio data generation.
[1049] System Program Overview
[1050] The program of this system is implemented using the following hardware and software.
[1051] Server: A computer system for storing, searching, generating, and providing data.
[1052] Database: Management of voice data, terms of use, user information, etc.
[1053] Website: The interface for users (voice actors and filmmakers) to access
[1054] Speech generation AI model: Artificial intelligence technology that analyzes input text and generates voice data
[1055] Voice data registration details
[1056] A user (voice actor) accesses the website and creates an account. When creating an account, the user enters the required information and clicks "Register," and the server saves the information in a database. Once registration is complete, the user (voice actor) logs in and registers their voice data, terms of use, and fees. Once the voice data is uploaded, the server validates it and saves it in a database.
[1057] More on finding and selecting voice actors
[1058] The user (filmmaker) logs in to the website and accesses the voice actor search page. Using the search function, they enter a keyword, such as "female female voice," and execute a search. The server searches the database for the relevant voice actor and presents the results to the user (filmmaker). The user can then check the detailed information from the displayed list of voice actors and select the one they need.
[1059] Voice data generation details
[1060] The user (video producer) enters text, such as lines, on the detailed information page of the voice actor they selected. When they click the "Generate" button, the server uses a voice generation AI model to convert the entered text into audio data. The generated audio data is then converted by the server into an appropriate format and provided to the user. The user (video producer) can then download this audio data and incorporate it into their own video work.
[1061] Specific examples
[1062] For example, if a user (filmmaker) is looking for a "young female voice" to use in their own animation work, they can proceed as follows: The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the search results. The user selects an appropriate voice actor from the list and enters text, such as "Hello, nice to meet you," on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed.
[1063] Example prompt sentence:
[1064] "Please convert the text "Hello, nice to meet you" into speech data in a young female voice."
[1065] In this way, the system of the present invention is a mechanism that realizes efficient and high-quality audio data generation and provides video producers with audio data that can be used quickly.
[1066] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1067] Step 1:
[1068] User (voice actor) account creation and information registration
[1069] The user (voice actor) accesses the website and enters the necessary information into the "Create Account" form, such as name, email address, and password.
[1070] When a user clicks "Register," the server receives the information and performs validation, such as checking the format of the email address and the strength of the password.
[1071] If validation passes, the server saves the account information in the database, allowing the user to log in.
[1072] After logging in, the user (voice actor) registers their own voice data, terms of use, and usage fees. After uploading the voice data, the server validates the data and stores it in the database. Specifically, it checks the format and size of the voice file.
[1073] Input: Account information entered by the user (name, email address, password) and voice data
[1074] Output: Account information and voice data stored in a database
[1075] Step 2:
[1076] User (video creator) login and voice actor search
[1077] The user (video creator) logs in to the website and accesses the voice actor search page. A username and password are required to log in.
[1078] The user (filmmaker) enters a keyword into the search bar, for example, "young woman's voice," and executes the search.
[1079] The server searches the database for a list of matching voice actors and retrieves the results, including the voice data and usage conditions that match the search criteria.
[1080] The server displays the search results to the user (video creator), which include the name of the voice actor, sample audio, terms of use, etc.
[1081] Input: The keyword that the user types into the search bar
[1082] Output: A list of voice actors displayed as search results
[1083] Step 3:
[1084] User (video creator) voice actor selection and text input
[1085] The user (filmmaker) selects the appropriate voice actor from the search results and accesses their details page.
[1086] The user checks the details and enters the text they want to convert into speech in the text entry field, for example, "Hello, nice to meet you."
[1087] The user clicks the "Generate" button, which sends the entered text to the server.
[1088] Input: The voice actor the user selects, and the text they enter into the text field.
[1089] Output: Text data sent to the server
[1090] Step 4:
[1091] Voice data generation by the server
[1092] The server sends the input text to a speech generation AI model, which analyzes the text and generates appropriate speech data.
[1093] The voice generation AI model generates speech based on the characteristics of a specified voice actor, for example, generating "Hello, nice to meet you" in a young female voice.
[1094] The server receives the generated audio data and performs format conversion and quality checks, including encoding the audio data and removing noise.
[1095] Input: Text data sent to the server
[1096] Output: Generated audio data
[1097] Step 5:
[1098] Providing audio data to users (video producers)
[1099] The server generates a download link for the generated audio data and provides the link to the user (video producer).
[1100] The user (video creator) downloads the audio data from the provided link and can incorporate the downloaded audio data into their own video work.
[1101] Input: Generated audio data
[1102] Output: Download link and audio data provided to the user
[1103] (Application example 1)
[1104] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1105] In existing content distribution systems, generating audio data and incorporating it in real time is extremely time-consuming, especially when providing an interactive experience. Furthermore, there is a lack of efficient means for audiovisual producers to search for and use the specific audio data they require. This creates a need for technology that can quickly and efficiently generate high-quality audio data and provide interactive experiences through visual display devices in real time.
[1106] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1107] In this invention, the server includes means for audio providers to register their own audio data and usage conditions, means for content creators to search for audio providers using keywords, means for converting input text into audio based on the audio data of the selected audio provider, means for incorporating the generated audio data into content in real time, and means for providing an interactive experience through a visual display device, thereby making it possible to quickly and efficiently generate audio data and provide an interactive experience through a visual display device in real time.
[1108] "Audio Data" means a digital file of sound recorded or generated by an audio provider.
[1109] "Audio provider" is a person or organization that provides audio data and sets the terms of use.
[1110] "Conditions of Use" are specific conditions and restrictions set by the audio provider when using audio data.
[1111] A "content creator" is a person or organization that creates content such as video or audio.
[1112] "Keywords" are specific words or phrases used in searches.
[1113] A "generative AI model" is an artificial intelligence technology that takes input text and converts it into voice data.
[1114] "Real time" refers to a state in which processing and responses occur almost simultaneously.
[1115] A "visual display device" is a device that allows a user to visually view content, and includes smartphones and head-mounted displays.
[1116] An "interactive experience" is one in which a user can interact with a system and receive responses in real time.
[1117] "Content" refers to digital media that includes information such as video and audio.
[1118] The present invention provides a system for generating audio data and incorporating it into content in real time, and is specifically implemented with the following configuration.
[1119] 1. System Overview
[1120] The system includes a voice provider, a content creator, a server, and a visual display device. The voice provider registers their voice data and terms of use, and the content creator searches for the voice provider using keywords, inputs text, and uses a voice generation AI model to generate voice data, which is then incorporated into the content in real time.
[1121] 2. Hardware and Software
[1122] Hardware: Smartphones, head-mounted displays, servers
[1123] Software: Web-based applications, databases (e.g., MySQL), speech generation AI models (e.g., Tacotron, WaveNet)
[1124] 3. System Processing Procedures
[1125] 1. Voice provider voice data registration
[1126] The voice provider visits the website and creates an account.
[1127] Once you enter your account information and register, the server stores that information in a database.
[1128] After completing the registration, the voice provider inputs his / her voice data, the terms of use, and the usage fee, and transmits them to the server.
[1129] The server validates the uploaded audio data and stores it in a database.
[1130] 2. Search and select an audio provider
[1131] A content creator logs into a website and visits the search page.
[1132] A content creator enters keywords and performs a search.
[1133] The server searches the database for the appropriate audio provider and presents the results to the content creator.
[1134] The content creator checks the detailed information from the displayed list of audio providers and selects one.
[1135] 3. Generating audio data
[1136] Enter text on the details page of the audio provider that the content creator has selected.
[1137] The server uses a generative AI model to convert the input text into audio data.
[1138] After successful speech generation, the server converts the generated speech data into an appropriate format and provides an interactive experience through a visual display device in real time.
[1139] 4. Specific Examples
[1140] For example, if a content creator searches for "young female voice" and enters the text "Hello, nice to meet you," the server will use the generative AI model to convert this text into speech data and incorporate the generated speech data into the content in real time, thereby providing an interactive experience to the user through a visual display device.
[1141] 5. Examples of prompts
[1142] Generative AI models:
[1143] Prompt: Convert the following text into the voice of the voice actor you provided: "Hello, nice to meet you."
[1144] Expected output: The text "Hello, nice to meet you" will be output as audio data from the selected young female voice actress.
[1145] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1146] Step 1:
[1147] The voice provider accesses the website and creates an account. The account information entered by the voice provider is sent to the server, which stores the information in a database.
[1148] Input: Account information input by the voice provider
[1149] Data processing: The server receives the account information and stores it in the database in the appropriate format.
[1150] Output: Account registration completion notification
[1151] Step 2:
[1152] After completing registration, the voice provider enters their voice data, terms of use, and usage fees into a form on the website and submits it to the server, which then validates the uploaded voice data and stores it in a database.
[1153] Input: Input of audio data, terms of use, and fees by audio provider
[1154] Data processing: The server receives the data, validates the voice data, and saves it to the database.
[1155] Output: Notification of completion of voice data registration
[1156] Step 3:
[1157] A content creator logs in to the website and accesses the search page. The content creator enters keywords and performs a search. The server searches the database for the corresponding audio provider and presents the results to the content creator.
[1158] Input: Keywords entered by content creators
[1159] Data processing: The server performs a keyword search and extracts the corresponding voice providers from the database.
[1160] Output: Presenting search results
[1161] Step 4:
[1162] The content creator checks the details of the audio providers from the list displayed and selects one. On the details page of the selected audio provider, they input text and request audio generation.
[1163] Input: Content creator selects audio provider and inputs text
[1164] Data processing: The server receives the text and invokes a generative AI model based on the voice data of the selected voice provider.
[1165] Output: Notification of execution of generation request
[1166] Step 5:
[1167] The server uses the generative AI model to convert the input text into speech data, and after successful speech generation, converts the generated speech data into an appropriate format and incorporates it into the content in real time.
[1168] Input: Text input to the generative AI model
[1169] Data processing: The generative AI model converts the text into audio data, and the server retrieves the audio data.
[1170] Output: Generated audio data
[1171] Step 6:
[1172] The generated audio data is incorporated into the content through a visual display device to provide an interactive experience for the user.
[1173] Input: Generated audio data
[1174] Data processing: Integrating audio data into the content of a visual display
[1175] Output: Interactive experience through a visual display device
[1176] This series of processing steps allows content creators to efficiently generate audio data and deliver real-time interactive experiences through visual display devices.
[1177] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1178] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user-created animation or test video. A specific embodiment of the system of the present invention will be described below.
[1179] System Overview
[1180] This system consists of four main components.
[1181] 1. A means for voice actors to register their own voice data and terms of use
[1182] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[1183] 3. Means for converting input text into speech based on the voice data of the selected voice actor.
[1184] 4. A means to recognize the user's emotions using an emotion engine and generate voice data based on that information
[1185] Program processing
[1186] 1. Voice actor voice data registration
[1187] A user (voice actor) accesses the website and creates an account.
[1188] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[1189] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[1190] The server validates the uploaded audio data and stores it in a database.
[1191] 2. Search and select a voice actor
[1192] The user (video creator) logs in to the website and accesses the voice actor search page.
[1193] The user (video creator) enters keywords and performs a search.
[1194] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[1195] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1196] 3. Emotion-Recognition Assistance
[1197] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[1198] The device recognizes the user's emotions in real time from their facial expressions and voice, and sends that information to the server.
[1199] The emotion engine assists in generating voice data containing appropriate intonation and emotional expression based on the recognized emotion.
[1200] 4. Generating Audio Data
[1201] The server uses a speech generation AI and emotion engine to convert the input text into voice data.
[1202] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[1203] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[1204] Specific examples
[1205] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. In this process, the emotion "joy" is recognized from the user's facial expressions and voice, and the server selects voice data from the voice actor based on that, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[1206] In this way, the system of the present invention can significantly improve the quality of video works by recognizing the user's emotions in real time and providing appropriate audio data, thereby significantly reducing production costs and time and providing an easy-to-use environment for users to easily create high-quality works.
[1207] The processing flow will be explained below.
[1208] Voice actor voice data registration process
[1209] Step 1:
[1210] The user (voice actor) accesses the website and opens the account creation page.
[1211] The device presents the user with an account creation form.
[1212] Step 2:
[1213] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[1214] The terminal sends the input information to the server.
[1215] Step 3:
[1216] The server creates a new user record in the database based on the received account information.
[1217] The server validates the input and saves it to the database.
[1218] Step 4:
[1219] The server sends a notification to the user that account registration is complete.
[1220] The server sends emails and site notifications.
[1221] Step 5:
[1222] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[1223] The device displays a voice registration form.
[1224] Step 6:
[1225] The user (voice actor) fills out the required information in the form and uploads the audio file.
[1226] The device sends voice data and other input information to the server.
[1227] Step 7:
[1228] The server validates the uploaded audio data and stores it in the database.
[1229] The server checks the format and size of the audio data and stores it in a database.
[1230] Step 8:
[1231] The server sends a notification to the user that the voice data has been registered.
[1232] The server sends emails and site notifications.
[1233] Voice Actor Search and Selection Process
[1234] Step 1:
[1235] The user (video creator) logs in to the website and accesses the voice actor search page.
[1236] The terminal displays a login form and, after successful authentication, a search page.
[1237] Step 2:
[1238] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[1239] The terminal sends the entered keyword to the server.
[1240] Step 3:
[1241] The server searches the database for voice actors that match the keywords.
[1242] The server executes a search query against the database and retrieves the results.
[1243] Step 4:
[1244] The server sends the search results to the user's (filmmaker's) device.
[1245] The server sends the search results in JSON format, which the device parses for display.
[1246] Step 5:
[1247] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1248] The device will display a detailed information page.
[1249] Assisted processes using emotion recognition
[1250] Step 1:
[1251] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[1252] The terminal displays an input text box.
[1253] Step 2:
[1254] The user (video creator) uses a camera and microphone to provide facial expressions and voice to the system.
[1255] The device captures the user's facial expressions and voice data in real time and sends it to the server.
[1256] Step 3:
[1257] The server uses an emotion engine to recognize the user's emotions in real time from the transmitted facial and voice data.
[1258] The server sends the input data to the emotion engine and analyzes the emotions.
[1259] Step 4:
[1260] Based on the recognized emotion, the server sends a request to the speech generation AI to generate speech data containing expressions that best fit the input text.
[1261] The server sends the emotion information and text to the speech generation AI.
[1262] Step 5:
[1263] The server receives the generated audio data, converts it into an appropriate format, and provides it to the user (video producer).
[1264] The server converts the audio data into MP3 or WAV format and sends it to the device.
[1265] Step 6:
[1266] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[1267] The device displays a download link for the audio data, which the user can then download and use.
[1268] Specific examples
[1269] For example, a user (video producer) searches for "a young female voice," selects a specific voice actor, and enters the text "Hello, nice to meet you." Next, the user uses a camera or microphone to provide the system with their own smiling face and cheerful voice, which the emotion engine recognizes as "joy." Based on this emotional information, the server sends a request to the voice generation AI, which generates voice data reflecting the emotion of "joy." The user can download this generated voice data and incorporate it into their own animated work, enabling them to create an emotionally rich production.
[1270] Example 2
[1271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1272] Conventionally, when generating audio data in video production, it has been difficult to provide audio that reflects the user's emotions. Furthermore, searching and managing voice data for voice actors is cumbersome, increasing production costs and time. This has led to the problem of making it difficult to efficiently produce high-quality video works.
[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1274] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for recognizing the user's emotions and generating voice data based on that information, and a means for incorporating the generated voice data into video. This makes it possible to efficiently generate high-quality voice data that reflects the user's emotions and incorporate it into video.
[1275] "Voice actors" are people who register their own voice data and play the role of providing audio for video production.
[1276] "Video producers" are people who generate the necessary audio data and perform editing work when producing a video work.
[1277] "Audio Data" means audio recordings provided by voice actors or audio generated by a generative AI model.
[1278] "Terms of Use" refers to the conditions and restrictions for using the voice data provided by the voice actors.
[1279] "Keywords" are words or phrases that filmmakers use to search for voice actors.
[1280] "Emotion recognition" is a technology that analyzes a user's facial expressions and voice and identifies their emotions in real time.
[1281] "Speech generation artificial intelligence" is a technology that generates natural-sounding speech based on input text and emotional data.
[1282] The "database" is an information collection system for efficiently managing and storing voice actor audio data and usage conditions.
[1283] "Text" refers to the string of characters entered by the video producer, and is the words or sentences that form the basis of the audio data.
[1284] "Video embedding" refers to the process of embedding the generated audio data as part of a video production.
[1285] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user's own animation or video work. A specific embodiment of the system of the present invention is described below.
[1286] This system consists of four main components.
[1287] 1. How to register voice data
[1288] A user (voice actor) accesses the website and creates an account by entering information such as name, email address, and password, and submitting it.
[1289] The server validates the entered information and saves it in the database. After successful saving, the user (voice actor) enters the voice data, terms of use, and fees, and uploads it.
[1290] The server validates the audio data and stores it in the database.
[1291] 2. How to search for voice actors
[1292] The user (video creator) logs in to the system and accesses the voice actor search page, where they enter keywords and perform a search.
[1293] The server searches the database for the relevant voice actor and displays the results as a list to the user (video producer).
[1294] The user (video creator) clicks and selects detailed information from the list of voice actors displayed.
[1295] 3. Text Input and Emotion Recognition
[1296] The user (video producer) goes to the details page of the voice actor they have chosen and enters the lines in the text input box.
[1297] The user's facial expressions and voice are captured in real time, and the device recognizes their emotions using emotion recognition software (e.g., OpenFace or Microsoft Emotion API).
[1298] The recognized emotion data is sent to a server.
[1299] 4. Audio data generation means
[1300] The server uses a speech generation AI (for example, Google Text-to-Speech API or Amazon Polly) to generate speech based on the input text and emotion data.
[1301] The server converts the generated audio data into an appropriate format and provides it to the user (video producer).
[1302] The user (video creator) downloads the generated audio data and incorporates it into their own video work.
[1303] Specific examples
[1304] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. During this process, the emotion "joy" is recognized from the user's facial expressions and voice. Based on this, the server selects voice data from the voice actor, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[1305] An example of a prompt to input to a generative AI model is as follows:
[1306] 1. "In a young female voice, please vocalize the lines "Hello" and "Nice to meet you" with a joyful emotion."
[1307] 2. "Can you say 'I'm looking forward to tomorrow!' in an energetic male voice with a surprised expression?"
[1308] 3. "Generate a recording of a middle-aged man quietly saying "goodbye" with sadness in his voice."
[1309] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[1310] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1311] Step 1: Create an account (a method for registering audio data)
[1312] Input: The user (voice actor) visits the website and enters their account information, including their name, email address, and password.
[1313] Data Processing: The server validates the entered information to ensure accuracy and format.
[1314] Output: If validation is successful, the server saves the account information to the database and sends the user a confirmation email.
[1315] Specific operation: The user (voice actor) enters the required information into the web form and clicks the "Create Account" button. The server saves the information in the database and automatically sends a confirmation email.
[1316] Step 2: Upload the audio data and terms of use
[1317] Input: After logging in, the user (voice actor) enters the audio data, terms of use, and usage fee into the web form and presses the upload button.
[1318] Data processing: The server receives the audio file and validates the file format and quality. It also checks the terms of use and fees.
[1319] Output: The voice data and terms of use that pass validation are saved in the database, and a success notification is displayed to the user.
[1320] Specific operation: The user (voice actor) selects an audio file, enters the required conditions, and clicks the "Upload" button. The server processes this and notifies the user of the result.
[1321] Step 3: Find a voice actor
[1322] Input: The user (video producer) logs into the system, accesses the voice actor search page, enters keywords, and presses the search button.
[1323] Data calculation: The server searches the database for the appropriate voice actor and filters the results that match the keywords.
[1324] Output: The search results are displayed as a list, and detailed information about the corresponding voice actors is presented to the user.
[1325] Specific operation: The user (filmmaker) enters keywords in the search box and clicks the "Search" button. The server retrieves the results and displays them on the page.
[1326] Step 4: Select a voice actor and enter text
[1327] Input: The user (video creator) clicks on the displayed voice actor details page and enters lines in the text input box.
[1328] Data processing: The server records the user's selection and temporarily stores the text data.
[1329] Output: A web page displays text entry boxes and other details that the user types and are sent to the server and stored.
[1330] Specific operation: The user (filmmaker) opens the details page, enters the dialogue text, and clicks the "Submit" button. The server receives and stores this information.
[1331] Step 5: Emotion Recognition
[1332] Input: While the user (video creator) is entering text, the device captures the user's facial expressions and voice in real time.
[1333] Data calculation: The device uses emotion recognition software (e.g., OpenFace or Microsoft Emotion API) to analyze the captured data and generate emotion data.
[1334] Output: The generated emotion data is sent to the server and associated with the text data.
[1335] How it works: While the user (filmmaker) is typing text, the device captures facial expressions and voice using the built-in camera and microphone, and runs emotion recognition software.
[1336] Step 6: Generate audio data
[1337] Input: The server sends the input text and emotion data to a speech generation AI (e.g., Google Text-to-Speech API or Amazon Polly).
[1338] Data calculation: The voice generation AI processes the input data based on the prompt sentence and generates voice data that reflects the emotion.
[1339] Output: The generated audio data is returned to the server and converted into the appropriate format.
[1340] Specific operation: The server passes text and emotion data to the speech generation AI, receives the generated speech data, and performs format conversion.
[1341] Step 7: Provide audio data
[1342] Input: The server provides the generated audio data to the user (video producer).
[1343] Data processing: The server rechecks the quality of the audio data and generates a download link.
[1344] Output: A page is displayed containing a link that allows the user (the filmmaker) to download the audio data.
[1345] Specific operation: The user (video creator) clicks on the download link, obtains the audio data, and incorporates it into the video work.
[1346] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[1347] (Application example 2)
[1348] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1349] A major challenge in current video production is the significant time and expense required to generate appropriate audio data. Particularly in the advertising field, emotive audio resonates with viewers and increases the effectiveness of advertising. However, existing speech synthesis technologies have limitations in expressing emotions, making it difficult to generate effective advertising audio. Therefore, there is a need for a system that can recognize users' actual emotions in real time and reflect them in audio data.
[1350] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for voice actors to register their own voice data and usage conditions, means for video producers to search for voice actors using keywords, means for converting input text into voice based on the voice data of the selected voice actor, means for recognizing the user's emotions and generating voice data based on those emotions, and means for incorporating the generated voice data into video. This makes it possible to generate emotionally rich voice data that reflects the user's emotions in real time.
[1351] The "voice actor registration means" is a means for a voice actor to register his / her voice data and usage conditions in the system.
[1352] The "voice actor search means" is a means for a video producer to use keywords to search for appropriate voice actors from a database.
[1353] The "voice conversion means" is a means for converting input text into voice based on the voice data of the selected voice actor.
[1354] The "emotion recognition means" is a means for recognizing the user's emotion and generating appropriate voice data based on that emotion.
[1355] The "voice generation means" is a means for converting input text into an emotionally rich voice based on the emotion recognized by the emotion recognition means.
[1356] The "video incorporation means" is a means for integrating the generated audio data into video content to generate a final video work.
[1357] The "database storage means" is a means for storing voice data registered by voice actors and the conditions for using the data in a database.
[1358] This invention relates to a system that recognizes a user's emotions in real time, generates voice data based on the emotions, and incorporates the voice data into video works and advertisements. This system includes a voice actor registration means, a voice actor search means, a voice conversion means, an emotion recognition means, a voice generation means, a video embedding means, and a database storage means.
[1359] composition
[1360] The system mainly consists of the following components:
[1361] 1. Hardware:
[1362] Smartphone or head-mounted display (HMD): Equipped with a camera and microphone to recognize the user's emotions.
[1363] Cloud database: A database (e.g., Google Firebase) for storing voice actor audio data and terms of use.
[1364] 2. Software:
[1365] Emotion recognition engine (Emotion AI): Recognizes emotions from the user's facial expressions and voice in real time.
[1366] Speech generation AI (Text-to-Speech engine): Converts input text into emotive speech.
[1367] Front-end framework (React Native): A mobile application to provide the user interface.
[1368] Program processing
[1369] The server first allows users to access the system and create an account. Voice actors register their voice data and terms of use, which are then stored in a database. This process is carried out through a front-end interface using React Native and is stored in Firebase.
[1370] The filmmaker (user) logs into the system using a smartphone or HMD and inputs specific keywords to search for suitable voice actors. Results are returned from the cloud database, and the user can select from a list of voice actors.
[1371] Based on the voice data of the selected voice actor, the user inputs a prompt, such as advertising text like "Buy now and get a special discount!". At this time, the emotion recognition engine (Emotion AI) analyzes the user's facial expressions and voice and collects emotional data in real time.
[1372] The emotion data collected by the emotion recognition engine is sent to the server, and then the voice generation AI converts the prompt sentence into voice based on this emotion data. For example, if the user expresses the emotion "excitement," voice data of the voice actor that reflects this emotion will be generated.
[1373] The generated audio data is integrated into the user's video content or advertisements using a video integration means, allowing the user to easily create high-quality video works that reflect emotional audio.
[1374] Specific examples
[1375] For example, if an advertiser uses a smartphone app to input the text "Buy now and get a special discount!", Emotion AI will recognize the emotion "excited" from the user's facial expressions and voice. The speech generation AI will then convert the text into speech, which will then be incorporated into the ad.
[1376] Prompt Sentence Examples
[1377] "Buy now and get a special discount!"
[1378] Recognized emotion: "Excitement"
[1379] Use voice data from a voice actor to convert this text into emotive speech.
[1380] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1381] Step 1:
[1382] A user accesses the system and creates an account. The user (voice actor) enters their voice data and terms of use and uploads it to the server. The server validates this information and stores it in a cloud database (e.g., Google Firebase). This creates a voice database that can be searched later.
[1383] Step 2:
[1384] The user (filmmaker) logs into the system and searches for a voice actor by entering a keyword. The entered keyword is sent to the server, which searches the database for a list of matching voice actors. The server displays the search results to the user, who then selects a voice actor from the displayed list.
[1385] Step 3:
[1386] The user (video creator) accesses the details page of the voice actor they selected and enters a prompt sentence (e.g., "Buy now and get a special discount!"). This entered text is sent to the server. At the same time, the emotion recognition engine (Emotion AI) collects emotional data from the user's facial expressions and voice in real time and sends it to the server.
[1387] Step 4:
[1388] The server sends prompts to a text-to-speech engine based on the acquired emotional data and the input text. The text-to-speech engine then generates appropriate, emotionally rich speech data. In this process, the emotional data sent from the emotion recognition engine is reflected in the speech generation.
[1389] Step 5:
[1390] The server converts the generated audio data into an appropriate format and provides it to the user's (video creator's) device. The user then downloads the generated audio data and integrates it into their own video content or advertisements. This results in the creation of a high-quality video work that includes richly expressive audio.
[1391] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1392] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1393] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1394] [Fourth embodiment]
[1395] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1396] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1397] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1398] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1399] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1400] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1401] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1402] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1403] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1404] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1405] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1406] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1407] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1408] The present invention relates to a system for easily generating audio data and incorporating it into a self-produced animation or test video. Specific embodiments of the system of the present invention will be described below.
[1409] System Overview
[1410] This system consists of three main components.
[1411] 1. A means for voice actors to register their own voice data and terms of use
[1412] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[1413] 3. Means for converting input text into speech based on the voice data of the selected voice actor and incorporating the generated speech data into the video.
[1414] Program processing
[1415] 1. Voice actor voice data registration
[1416] A user (voice actor) accesses the website and creates an account.
[1417] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[1418] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[1419] The server validates the uploaded audio data and stores it in a database.
[1420] 2. Search and select a voice actor
[1421] The user (video creator) logs in to the website and accesses the voice actor search page.
[1422] The user (video creator) enters keywords and performs a search.
[1423] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[1424] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1425] 3. Generating audio data
[1426] Enter text on the detailed information page of the voice actor selected by the user (video producer).
[1427] The server uses speech generation AI to convert the input text into audio data.
[1428] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[1429] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[1430] Specific examples
[1431] For example, consider the case where a user (filmmaker) is looking for a "young female voice" to use in one of their own animation works. The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the results. The user selects an appropriate voice actor from the results and enters the text "Hello, nice to meet you" on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed easily.
[1432] In this way, the system of the present invention provides an environment in which users can quickly and efficiently generate audio data and use it in video productions, significantly reducing production costs and time and enabling even amateur producers to achieve professional-looking results.
[1433] The processing flow will be explained below.
[1434] Voice actor voice data registration process
[1435] Step 1:
[1436] The user (voice actor) accesses the website and opens the account creation page.
[1437] The device presents the user with an account creation form.
[1438] Step 2:
[1439] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[1440] The terminal sends the input information to the server.
[1441] Step 3:
[1442] The server creates a new user record in the database based on the received account information.
[1443] The server validates the input and saves it to the database.
[1444] Step 4:
[1445] The server sends a notification to the user that account registration is complete.
[1446] The server sends emails and site notifications.
[1447] Step 5:
[1448] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[1449] The device displays a voice registration form.
[1450] Step 6:
[1451] The user (voice actor) fills out the required information in the form and uploads the audio file.
[1452] The device sends voice data and other input information to the server.
[1453] Step 7:
[1454] The server validates the uploaded audio data and stores it in the database.
[1455] The server checks the format and size of the audio data and stores it in a database.
[1456] Step 8:
[1457] The server sends a notification to the user that the voice data has been registered.
[1458] The server sends emails and site notifications.
[1459] Voice Actor Search and Selection Process
[1460] Step 1:
[1461] The user (video creator) logs in to the website and accesses the voice actor search page.
[1462] The terminal displays a login form and, after successful authentication, a search page.
[1463] Step 2:
[1464] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[1465] The terminal sends the entered keyword to the server.
[1466] Step 3:
[1467] The server searches the database for voice actors that match the keywords.
[1468] The server executes a search query against the database and retrieves the results.
[1469] Step 4:
[1470] The server sends the search results to the user's (filmmaker's) device.
[1471] The server sends the search results in JSON format, which the device parses for display.
[1472] Step 5:
[1473] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1474] The device will display a detailed information page.
[1475] Voice data generation process
[1476] Step 1:
[1477] On the detailed information page of the voice actor selected by the user (video producer), enter a line (e.g., "Hello, nice to meet you") in the text input box and click the "Generate Voice" button.
[1478] The device sends the entered text and the selected voice actor's ID to the server.
[1479] Step 2:
[1480] The server sends a voice generation request to the voice generation AI.
[1481] The server sends the text and voice actor ID to the voice generation API and receives the voice data.
[1482] Step 3:
[1483] The server converts the generated audio data into a specific format and sends it back to the user (video producer).
[1484] The server converts the audio data into MP3 or WAV format and sends it to the device.
[1485] Step 4:
[1486] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[1487] The device displays a download link for the audio data, which the user can then download and use.
[1488] Specific examples
[1489] For example, a user (video creator) searches for "the voice of a young woman," selects a specific voice actor, enters the text "Hello, nice to meet you," and generates the voice. Through this process, a series of steps are realized in which the generated voice data is received from the server, downloaded, and incorporated into the user's own animated work.
[1490] Example 1
[1491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1492] In today's video production, generating and incorporating audio data requires a significant amount of time and cost. Searching, selecting, and managing audio from professional voice actors is particularly complex, and the lack of an efficient system is a problem. This makes amateur productions and short-term projects difficult.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1494] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for incorporating the generated voice data into a video work, a means for converting input text into voice using a voice generation AI model, and a means for providing the generated voice data to users, thereby enabling video producers to efficiently generate and incorporate voice data.
[1495] A "voice actor" is an individual or group that uses their own voice to provide audio data.
[1496] "Audio data" refers to audio information provided by voice actors stored in digital data format.
[1497] "Conditions of use" refers to restrictions and permissions regarding the use of audio data, such as whether commercial use is permitted or not, and the expiration date of use.
[1498] "Filmmaker" means an individual or organization that produces a film work.
[1499] A "keyword" is a character string or phrase that a user enters to search for a voice actor.
[1500] A "voice generation AI model" is a type of artificial intelligence technology that analyzes input text and generates it as voice data.
[1501] "Text" refers to the written information that is the basis for generating speech.
[1502] A "server" is a computer system that processes various types of data and provides services to users.
[1503] A "database" is a system that systematically stores and manages information such as voice data and usage conditions.
[1504] "Prompt sentence" refers to an example of text that a user inputs into a speech-generation AI model.
[1505] This invention relates to a system that allows filmmakers to efficiently generate audio data and incorporate it into their video works. This system has the means to register, search, generate, and provide audio data, and by using a generative AI model, it realizes rapid and high-quality audio data generation.
[1506] System Program Overview
[1507] The program of this system is implemented using the following hardware and software.
[1508] Server: A computer system for storing, searching, generating, and providing data.
[1509] Database: Management of voice data, terms of use, user information, etc.
[1510] Website: The interface for users (voice actors and filmmakers) to access
[1511] Speech generation AI model: Artificial intelligence technology that analyzes input text and generates voice data
[1512] Voice data registration details
[1513] A user (voice actor) accesses the website and creates an account. When creating an account, the user enters the required information and clicks "Register," and the server saves the information in a database. Once registration is complete, the user (voice actor) logs in and registers their voice data, terms of use, and fees. Once the voice data is uploaded, the server validates it and saves it in a database.
[1514] More on finding and selecting voice actors
[1515] The user (filmmaker) logs in to the website and accesses the voice actor search page. Using the search function, they enter a keyword, such as "female female voice," and execute a search. The server searches the database for the relevant voice actor and presents the results to the user (filmmaker). The user can then check the detailed information from the displayed list of voice actors and select the one they need.
[1516] Voice data generation details
[1517] The user (video producer) enters text, such as lines, on the detailed information page of the voice actor they selected. When they click the "Generate" button, the server uses a voice generation AI model to convert the entered text into audio data. The generated audio data is then converted by the server into an appropriate format and provided to the user. The user (video producer) can then download this audio data and incorporate it into their own video work.
[1518] Specific examples
[1519] For example, if a user (filmmaker) is looking for a "young female voice" to use in their own animation work, they can proceed as follows: The user first logs into the system and enters "young female voice" on the search page. The server searches the database for matching voice actors and displays the search results. The user selects an appropriate voice actor from the list and enters text, such as "Hello, nice to meet you," on the actor's details page to request voice generation. The server then uses voice generation AI to convert this text into voice and provides the generated voice data to the user. By incorporating this voice data into the animation, a high-quality work can be completed.
[1520] Example prompt sentence:
[1521] "Please convert the text "Hello, nice to meet you" into speech data in a young female voice."
[1522] In this way, the system of the present invention is a mechanism that realizes efficient and high-quality audio data generation and provides video producers with audio data that can be used quickly.
[1523] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1524] Step 1:
[1525] User (voice actor) account creation and information registration
[1526] The user (voice actor) accesses the website and enters the necessary information into the "Create Account" form, such as name, email address, and password.
[1527] When a user clicks "Register," the server receives the information and performs validation, such as checking the format of the email address and the strength of the password.
[1528] If validation passes, the server saves the account information in the database, allowing the user to log in.
[1529] After logging in, the user (voice actor) registers their own voice data, terms of use, and usage fees. After uploading the voice data, the server validates the data and stores it in the database. Specifically, it checks the format and size of the voice file.
[1530] Input: Account information entered by the user (name, email address, password) and voice data
[1531] Output: Account information and voice data stored in a database
[1532] Step 2:
[1533] User (video creator) login and voice actor search
[1534] The user (video creator) logs in to the website and accesses the voice actor search page. A username and password are required to log in.
[1535] The user (filmmaker) enters a keyword into the search bar, for example, "young woman's voice," and executes the search.
[1536] The server searches the database for a list of matching voice actors and retrieves the results, including the voice data and usage conditions that match the search criteria.
[1537] The server displays the search results to the user (video creator), which include the name of the voice actor, sample audio, terms of use, etc.
[1538] Input: The keyword that the user types into the search bar
[1539] Output: A list of voice actors displayed as search results
[1540] Step 3:
[1541] User (video creator) voice actor selection and text input
[1542] The user (filmmaker) selects the appropriate voice actor from the search results and accesses their details page.
[1543] The user checks the details and enters the text they want to convert into speech in the text entry field, for example, "Hello, nice to meet you."
[1544] The user clicks the "Generate" button, which sends the entered text to the server.
[1545] Input: The voice actor the user selects, and the text they enter into the text field.
[1546] Output: Text data sent to the server
[1547] Step 4:
[1548] Voice data generation by the server
[1549] The server sends the input text to a speech generation AI model, which analyzes the text and generates appropriate speech data.
[1550] The voice generation AI model generates speech based on the characteristics of a specified voice actor, for example, generating "Hello, nice to meet you" in a young female voice.
[1551] The server receives the generated audio data and performs format conversion and quality checks, including encoding the audio data and removing noise.
[1552] Input: Text data sent to the server
[1553] Output: Generated audio data
[1554] Step 5:
[1555] Providing audio data to users (video producers)
[1556] The server generates a download link for the generated audio data and provides the link to the user (video producer).
[1557] The user (video creator) downloads the audio data from the provided link and can incorporate the downloaded audio data into their own video work.
[1558] Input: Generated audio data
[1559] Output: Download link and audio data provided to the user
[1560] (Application example 1)
[1561] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1562] In existing content distribution systems, generating audio data and incorporating it in real time is extremely time-consuming, especially when providing an interactive experience. Furthermore, there is a lack of efficient means for audiovisual producers to search for and use the specific audio data they require. This creates a need for technology that can quickly and efficiently generate high-quality audio data and provide interactive experiences through visual display devices in real time.
[1563] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1564] In this invention, the server includes means for audio providers to register their own audio data and usage conditions, means for content creators to search for audio providers using keywords, means for converting input text into audio based on the audio data of the selected audio provider, means for incorporating the generated audio data into content in real time, and means for providing an interactive experience through a visual display device, thereby making it possible to quickly and efficiently generate audio data and provide an interactive experience through a visual display device in real time.
[1565] "Audio Data" means a digital file of sound recorded or generated by an audio provider.
[1566] "Audio provider" is a person or organization that provides audio data and sets the terms of use.
[1567] "Conditions of Use" are specific conditions and restrictions set by the audio provider when using audio data.
[1568] A "content creator" is a person or organization that creates content such as video or audio.
[1569] "Keywords" are specific words or phrases used in searches.
[1570] A "generative AI model" is an artificial intelligence technology that takes input text and converts it into voice data.
[1571] "Real time" refers to a state in which processing and responses occur almost simultaneously.
[1572] A "visual display device" is a device that allows a user to visually view content, and includes smartphones and head-mounted displays.
[1573] An "interactive experience" is one in which a user can interact with a system and receive responses in real time.
[1574] "Content" refers to digital media that includes information such as video and audio.
[1575] The present invention provides a system for generating audio data and incorporating it into content in real time, and is specifically implemented with the following configuration.
[1576] 1. System Overview
[1577] The system includes a voice provider, a content creator, a server, and a visual display device. The voice provider registers their voice data and terms of use, and the content creator searches for the voice provider using keywords, inputs text, and uses a voice generation AI model to generate voice data, which is then incorporated into the content in real time.
[1578] 2. Hardware and Software
[1579] Hardware: Smartphones, head-mounted displays, servers
[1580] Software: Web-based applications, databases (e.g., MySQL), speech generation AI models (e.g., Tacotron, WaveNet)
[1581] 3. System Processing Procedures
[1582] 1. Voice provider voice data registration
[1583] The voice provider visits the website and creates an account.
[1584] Once you enter your account information and register, the server stores that information in a database.
[1585] After completing the registration, the voice provider inputs his / her voice data, the terms of use, and the usage fee, and transmits them to the server.
[1586] The server validates the uploaded audio data and stores it in a database.
[1587] 2. Search and select an audio provider
[1588] A content creator logs into a website and visits the search page.
[1589] A content creator enters keywords and performs a search.
[1590] The server searches the database for the appropriate audio provider and presents the results to the content creator.
[1591] The content creator checks the detailed information from the displayed list of audio providers and selects one.
[1592] 3. Generating audio data
[1593] Enter text on the details page of the audio provider that the content creator has selected.
[1594] The server uses a generative AI model to convert the input text into audio data.
[1595] After successful speech generation, the server converts the generated speech data into an appropriate format and provides an interactive experience through a visual display device in real time.
[1596] 4. Specific Examples
[1597] For example, if a content creator searches for "young female voice" and enters the text "Hello, nice to meet you," the server will use the generative AI model to convert this text into speech data and incorporate the generated speech data into the content in real time, thereby providing an interactive experience to the user through a visual display device.
[1598] 5. Examples of prompts
[1599] Generative AI models:
[1600] Prompt: Convert the following text into the voice of the voice actor you provided: "Hello, nice to meet you."
[1601] Expected output: The text "Hello, nice to meet you" will be output as audio data from the selected young female voice actress.
[1602] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1603] Step 1:
[1604] The voice provider accesses the website and creates an account. The account information entered by the voice provider is sent to the server, which stores the information in a database.
[1605] Input: Account information input by the voice provider
[1606] Data processing: The server receives the account information and stores it in the database in the appropriate format.
[1607] Output: Account registration completion notification
[1608] Step 2:
[1609] After completing registration, the voice provider enters their voice data, terms of use, and usage fees into a form on the website and submits it to the server, which then validates the uploaded voice data and stores it in a database.
[1610] Input: Input of audio data, terms of use, and fees by audio provider
[1611] Data processing: The server receives the data, validates the voice data, and saves it to the database.
[1612] Output: Notification of completion of voice data registration
[1613] Step 3:
[1614] A content creator logs in to the website and accesses the search page. The content creator enters keywords and performs a search. The server searches the database for the corresponding audio provider and presents the results to the content creator.
[1615] Input: Keywords entered by content creators
[1616] Data processing: The server performs a keyword search and extracts the corresponding voice providers from the database.
[1617] Output: Presenting search results
[1618] Step 4:
[1619] The content creator checks the details of the audio providers from the list displayed and selects one. On the details page of the selected audio provider, they input text and request audio generation.
[1620] Input: Content creator selects audio provider and inputs text
[1621] Data processing: The server receives the text and invokes a generative AI model based on the voice data of the selected voice provider.
[1622] Output: Notification of execution of generation request
[1623] Step 5:
[1624] The server uses the generative AI model to convert the input text into speech data, and after successful speech generation, converts the generated speech data into an appropriate format and incorporates it into the content in real time.
[1625] Input: Text input to the generative AI model
[1626] Data processing: The generative AI model converts the text into audio data, and the server retrieves the audio data.
[1627] Output: Generated audio data
[1628] Step 6:
[1629] The generated audio data is incorporated into the content through a visual display device to provide an interactive experience for the user.
[1630] Input: Generated audio data
[1631] Data processing: Integrating audio data into the content of a visual display
[1632] Output: Interactive experience through a visual display device
[1633] This series of processing steps allows content creators to efficiently generate audio data and deliver real-time interactive experiences through visual display devices.
[1634] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1635] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user-created animation or test video. A specific embodiment of the system of the present invention will be described below.
[1636] System Overview
[1637] This system consists of four main components.
[1638] 1. A means for voice actors to register their own voice data and terms of use
[1639] 2. A way for filmmakers to search for voice actors using keywords and obtain their audio data.
[1640] 3. Means for converting input text into speech based on the voice data of the selected voice actor.
[1641] 4. A means to recognize the user's emotions using an emotion engine and generate voice data based on that information
[1642] Program processing
[1643] 1. Voice actor voice data registration
[1644] A user (voice actor) accesses the website and creates an account.
[1645] When a user (voice actor) registers by entering their account information, the server stores that information in a database.
[1646] After completing the registration, the user (voice actor) further inputs his / her voice data, terms of use, and usage fee, and transmits them to the server.
[1647] The server validates the uploaded audio data and stores it in a database.
[1648] 2. Search and select a voice actor
[1649] The user (video creator) logs in to the website and accesses the voice actor search page.
[1650] The user (video creator) enters keywords and performs a search.
[1651] The server searches the database for the appropriate voice actor and presents the results to the user (filmmaker).
[1652] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1653] 3. Emotion-Recognition Assistance
[1654] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[1655] The device recognizes the user's emotions in real time from their facial expressions and voice, and sends that information to the server.
[1656] The emotion engine assists in generating voice data containing appropriate intonation and emotional expression based on the recognized emotion.
[1657] 4. Generating Audio Data
[1658] The server uses a speech generation AI and emotion engine to convert the input text into voice data.
[1659] After the audio is successfully generated, the server converts the generated audio data into an appropriate format and provides it to the user (video creator).
[1660] Users (filmmakers) can download the generated audio data and incorporate it into their own video works.
[1661] Specific examples
[1662] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. In this process, the emotion "joy" is recognized from the user's facial expressions and voice, and the server selects voice data from the voice actor based on that, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[1663] In this way, the system of the present invention can significantly improve the quality of video works by recognizing the user's emotions in real time and providing appropriate audio data, thereby significantly reducing production costs and time and providing an easy-to-use environment for users to easily create high-quality works.
[1664] The processing flow will be explained below.
[1665] Voice actor voice data registration process
[1666] Step 1:
[1667] The user (voice actor) accesses the website and opens the account creation page.
[1668] The device presents the user with an account creation form.
[1669] Step 2:
[1670] The user (voice actor) enters account information (name, email address, password, etc.) and clicks the "Register" button.
[1671] The terminal sends the input information to the server.
[1672] Step 3:
[1673] The server creates a new user record in the database based on the received account information.
[1674] The server validates the input and saves it to the database.
[1675] Step 4:
[1676] The server sends a notification to the user that account registration is complete.
[1677] The server sends emails and site notifications.
[1678] Step 5:
[1679] The user (voice actor) accesses a form to register his / her voice data, terms of use, fees, etc.
[1680] The device displays a voice registration form.
[1681] Step 6:
[1682] The user (voice actor) fills out the required information in the form and uploads the audio file.
[1683] The device sends voice data and other input information to the server.
[1684] Step 7:
[1685] The server validates the uploaded audio data and stores it in the database.
[1686] The server checks the format and size of the audio data and stores it in a database.
[1687] Step 8:
[1688] The server sends a notification to the user that the voice data has been registered.
[1689] The server sends emails and site notifications.
[1690] Voice Actor Search and Selection Process
[1691] Step 1:
[1692] The user (video creator) logs in to the website and accesses the voice actor search page.
[1693] The terminal displays a login form and, after successful authentication, a search page.
[1694] Step 2:
[1695] The user (video producer) enters a keyword (e.g., "the voice of a young woman") and clicks the "Search" button.
[1696] The terminal sends the entered keyword to the server.
[1697] Step 3:
[1698] The server searches the database for voice actors that match the keywords.
[1699] The server executes a search query against the database and retrieves the results.
[1700] Step 4:
[1701] The server sends the search results to the user's (filmmaker's) device.
[1702] The server sends the search results in JSON format, which the device parses for display.
[1703] Step 5:
[1704] The user (video producer) checks the detailed information from the displayed list of voice actors and selects one.
[1705] The device will display a detailed information page.
[1706] Assisted processes using emotion recognition
[1707] Step 1:
[1708] On the detailed information page of the voice actor selected by the user (video producer), enter the lines in the text input box.
[1709] The terminal displays an input text box.
[1710] Step 2:
[1711] The user (video creator) uses a camera and microphone to provide facial expressions and voice to the system.
[1712] The device captures the user's facial expressions and voice data in real time and sends it to the server.
[1713] Step 3:
[1714] The server uses an emotion engine to recognize the user's emotions in real time from the transmitted facial and voice data.
[1715] The server sends the input data to the emotion engine and analyzes the emotions.
[1716] Step 4:
[1717] Based on the recognized emotion, the server sends a request to the speech generation AI to generate speech data containing expressions that best fit the input text.
[1718] The server sends the emotion information and text to the speech generation AI.
[1719] Step 5:
[1720] The server receives the generated audio data, converts it into an appropriate format, and provides it to the user (video producer).
[1721] The server converts the audio data into MP3 or WAV format and sends it to the device.
[1722] Step 6:
[1723] The user (video producer) downloads the received audio data and incorporates it into their own video work.
[1724] The device displays a download link for the audio data, which the user can then download and use.
[1725] Specific examples
[1726] For example, a user (video producer) searches for "a young female voice," selects a specific voice actor, and enters the text "Hello, nice to meet you." Next, the user uses a camera or microphone to provide the system with their own smiling face and cheerful voice, which the emotion engine recognizes as "joy." Based on this emotional information, the server sends a request to the voice generation AI, which generates voice data reflecting the emotion of "joy." The user can download this generated voice data and incorporate it into their own animated work, enabling them to create an emotionally rich production.
[1727] Example 2
[1728] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1729] Conventionally, when generating audio data in video production, it has been difficult to provide audio that reflects the user's emotions. Furthermore, searching and managing voice data for voice actors is cumbersome, increasing production costs and time. This has led to the problem of making it difficult to efficiently produce high-quality video works.
[1730] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1731] In this invention, the server includes a means for voice actors to register their own voice data and usage conditions, a means for video producers to search for voice actors using keywords, a means for converting input text into voice based on the voice data of the selected voice actor, a means for recognizing the user's emotions and generating voice data based on that information, and a means for incorporating the generated voice data into video. This makes it possible to efficiently generate high-quality voice data that reflects the user's emotions and incorporate it into video.
[1732] "Voice actors" are people who register their own voice data and play the role of providing audio for video production.
[1733] "Video producers" are people who generate the necessary audio data and perform editing work when producing a video work.
[1734] "Audio Data" means audio recordings provided by voice actors or audio generated by a generative AI model.
[1735] "Terms of Use" refers to the conditions and restrictions for using the voice data provided by the voice actors.
[1736] "Keywords" are words or phrases that filmmakers use to search for voice actors.
[1737] "Emotion recognition" is a technology that analyzes a user's facial expressions and voice and identifies their emotions in real time.
[1738] "Speech generation artificial intelligence" is a technology that generates natural-sounding speech based on input text and emotional data.
[1739] The "database" is an information collection system for efficiently managing and storing voice actor audio data and usage conditions.
[1740] "Text" refers to the string of characters entered by the video producer, and is the words or sentences that form the basis of the audio data.
[1741] "Video embedding" refers to the process of embedding the generated audio data as part of a video production.
[1742] The present invention relates to a system that recognizes a user's emotions and provides appropriate audio data when generating audio data and incorporating it into a user's own animation or video work. A specific embodiment of the system of the present invention is described below.
[1743] This system consists of four main components.
[1744] 1. How to register voice data
[1745] A user (voice actor) accesses the website and creates an account by entering information such as name, email address, and password, and submitting it.
[1746] The server validates the entered information and saves it in the database. After successful saving, the user (voice actor) enters the voice data, terms of use, and fees, and uploads it.
[1747] The server validates the audio data and stores it in the database.
[1748] 2. How to search for voice actors
[1749] The user (video creator) logs in to the system and accesses the voice actor search page, where they enter keywords and perform a search.
[1750] The server searches the database for the relevant voice actor and displays the results as a list to the user (video producer).
[1751] The user (video creator) clicks and selects detailed information from the list of voice actors displayed.
[1752] 3. Text Input and Emotion Recognition
[1753] The user (video producer) goes to the details page of the voice actor they have chosen and enters the lines in the text input box.
[1754] The user's facial expressions and voice are captured in real time, and the device recognizes their emotions using emotion recognition software (e.g., OpenFace or Microsoft Emotion API).
[1755] The recognized emotion data is sent to a server.
[1756] 4. Audio data generation means
[1757] The server uses a speech generation AI (for example, Google Text-to-Speech API or Amazon Polly) to generate speech based on the input text and emotion data.
[1758] The server converts the generated audio data into an appropriate format and provides it to the user (video producer).
[1759] The user (video creator) downloads the generated audio data and incorporates it into their own video work.
[1760] Specific examples
[1761] For example, a user (video creator) searches for "a young female voice," selects a specific voice actor, enters the text "Hello, nice to meet you," and receives assistance from the emotion engine. During this process, the emotion "joy" is recognized from the user's facial expressions and voice. Based on this, the server selects voice data from the voice actor, and the generated voice data reflects that emotion. This results in realistic, emotionally rich voice data that users can easily incorporate into their own animations and video works.
[1762] An example of a prompt to input to a generative AI model is as follows:
[1763] 1. "In a young female voice, please vocalize the lines "Hello" and "Nice to meet you" with a joyful emotion."
[1764] 2. "Can you say 'I'm looking forward to tomorrow!' in an energetic male voice with a surprised expression?"
[1765] 3. "Generate a recording of a middle-aged man quietly saying "goodbye" with sadness in his voice."
[1766] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[1767] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1768] Step 1: Create an account (a method for registering audio data)
[1769] Input: The user (voice actor) visits the website and enters their account information, including their name, email address, and password.
[1770] Data Processing: The server validates the entered information to ensure accuracy and format.
[1771] Output: If validation is successful, the server saves the account information to the database and sends the user a confirmation email.
[1772] Specific operation: The user (voice actor) enters the required information into the web form and clicks the "Create Account" button. The server saves the information in the database and automatically sends a confirmation email.
[1773] Step 2: Upload the audio data and terms of use
[1774] Input: After logging in, the user (voice actor) enters the audio data, terms of use, and usage fee into the web form and presses the upload button.
[1775] Data processing: The server receives the audio file and validates the file format and quality. It also checks the terms of use and fees.
[1776] Output: The voice data and terms of use that pass validation are saved in the database, and a success notification is displayed to the user.
[1777] Specific operation: The user (voice actor) selects an audio file, enters the required conditions, and clicks the "Upload" button. The server processes this and notifies the user of the result.
[1778] Step 3: Find a voice actor
[1779] Input: The user (video producer) logs into the system, accesses the voice actor search page, enters keywords, and presses the search button.
[1780] Data calculation: The server searches the database for the appropriate voice actor and filters the results that match the keywords.
[1781] Output: The search results are displayed as a list, and detailed information about the corresponding voice actors is presented to the user.
[1782] Specific operation: The user (filmmaker) enters keywords in the search box and clicks the "Search" button. The server retrieves the results and displays them on the page.
[1783] Step 4: Select a voice actor and enter text
[1784] Input: The user (video creator) clicks on the displayed voice actor details page and enters lines in the text input box.
[1785] Data processing: The server records the user's selection and temporarily stores the text data.
[1786] Output: A web page displays text entry boxes and other details that the user types and are sent to the server and stored.
[1787] Specific operation: The user (filmmaker) opens the details page, enters the dialogue text, and clicks the "Submit" button. The server receives and stores this information.
[1788] Step 5: Emotion Recognition
[1789] Input: While the user (video creator) is entering text, the device captures the user's facial expressions and voice in real time.
[1790] Data calculation: The device uses emotion recognition software (e.g., OpenFace or Microsoft Emotion API) to analyze the captured data and generate emotion data.
[1791] Output: The generated emotion data is sent to the server and associated with the text data.
[1792] How it works: While the user (filmmaker) is typing text, the device captures facial expressions and voice using the built-in camera and microphone, and runs emotion recognition software.
[1793] Step 6: Generate audio data
[1794] Input: The server sends the input text and emotion data to a speech generation AI (e.g., Google Text-to-Speech API or Amazon Polly).
[1795] Data calculation: The voice generation AI processes the input data based on the prompt sentence and generates voice data that reflects the emotion.
[1796] Output: The generated audio data is returned to the server and converted into the appropriate format.
[1797] Specific operation: The server passes text and emotion data to the speech generation AI, receives the generated speech data, and performs format conversion.
[1798] Step 7: Provide audio data
[1799] Input: The server provides the generated audio data to the user (video producer).
[1800] Data processing: The server rechecks the quality of the audio data and generates a download link.
[1801] Output: A page is displayed containing a link that allows the user (the filmmaker) to download the audio data.
[1802] Specific operation: The user (video creator) clicks on the download link, obtains the audio data, and incorporates it into the video work.
[1803] This makes it possible to efficiently generate high-quality audio data that reflects the user's emotions and incorporate it into video.
[1804] (Application example 2)
[1805] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1806] A major challenge in current video production is the significant time and expense required to generate appropriate audio data. Particularly in the advertising field, emotive audio resonates with viewers and increases the effectiveness of advertising. However, existing speech synthesis technologies have limitations in expressing emotions, making it difficult to generate effective advertising audio. Therefore, there is a need for a system that can recognize users' actual emotions in real time and reflect them in audio data.
[1807] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for voice actors to register their own voice data and usage conditions, means for video producers to search for voice actors using keywords, means for converting input text into voice based on the voice data of the selected voice actor, means for recognizing the user's emotions and generating voice data based on those emotions, and means for incorporating the generated voice data into video. This makes it possible to generate emotionally rich voice data that reflects the user's emotions in real time.
[1808] The "voice actor registration means" is a means for a voice actor to register his / her voice data and usage conditions in the system.
[1809] The "voice actor search means" is a means for a video producer to use keywords to search for appropriate voice actors from a database.
[1810] The "voice conversion means" is a means for converting input text into voice based on the voice data of the selected voice actor.
[1811] The "emotion recognition means" is a means for recognizing the user's emotion and generating appropriate voice data based on that emotion.
[1812] The "voice generation means" is a means for converting input text into an emotionally rich voice based on the emotion recognized by the emotion recognition means.
[1813] The "video incorporation means" is a means for integrating the generated audio data into video content to generate a final video work.
[1814] The "database storage means" is a means for storing voice data registered by voice actors and the conditions for using the data in a database.
[1815] This invention relates to a system that recognizes a user's emotions in real time, generates voice data based on the emotions, and incorporates the voice data into video works and advertisements. This system includes a voice actor registration means, a voice actor search means, a voice conversion means, an emotion recognition means, a voice generation means, a video embedding means, and a database storage means.
[1816] composition
[1817] The system mainly consists of the following components:
[1818] 1. Hardware:
[1819] Smartphone or head-mounted display (HMD): Equipped with a camera and microphone to recognize the user's emotions.
[1820] Cloud database: A database (e.g., Google Firebase) for storing voice actor audio data and terms of use.
[1821] 2. Software:
[1822] Emotion recognition engine (Emotion AI): Recognizes emotions from the user's facial expressions and voice in real time.
[1823] Speech generation AI (Text-to-Speech engine): Converts input text into emotive speech.
[1824] Front-end framework (React Native): A mobile application to provide the user interface.
[1825] Program processing
[1826] The server first allows users to access the system and create an account. Voice actors register their voice data and terms of use, which are then stored in a database. This process is carried out through a front-end interface using React Native and is stored in Firebase.
[1827] The filmmaker (user) logs into the system using a smartphone or HMD and inputs specific keywords to search for suitable voice actors. Results are returned from the cloud database, and the user can select from a list of voice actors.
[1828] Based on the voice data of the selected voice actor, the user inputs a prompt, such as advertising text like "Buy now and get a special discount!". At this time, the emotion recognition engine (Emotion AI) analyzes the user's facial expressions and voice and collects emotional data in real time.
[1829] The emotion data collected by the emotion recognition engine is sent to the server, and then the voice generation AI converts the prompt sentence into voice based on this emotion data. For example, if the user expresses the emotion "excitement," voice data of the voice actor that reflects this emotion will be generated.
[1830] The generated audio data is integrated into the user's video content or advertisements using a video integration means, allowing the user to easily create high-quality video works that reflect emotional audio.
[1831] Specific examples
[1832] For example, if an advertiser uses a smartphone app to input the text "Buy now and get a special discount!", Emotion AI will recognize the emotion "excited" from the user's facial expressions and voice. The speech generation AI will then convert the text into speech, which will then be incorporated into the ad.
[1833] Prompt Sentence Examples
[1834] "Buy now and get a special discount!"
[1835] Recognized emotion: "Excitement"
[1836] Use voice data from a voice actor to convert this text into emotive speech.
[1837] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1838] Step 1:
[1839] A user accesses the system and creates an account. The user (voice actor) enters their voice data and terms of use and uploads it to the server. The server validates this information and stores it in a cloud database (e.g., Google Firebase). This creates a voice database that can be searched later.
[1840] Step 2:
[1841] The user (filmmaker) logs into the system and searches for a voice actor by entering a keyword. The entered keyword is sent to the server, which searches the database for a list of matching voice actors. The server displays the search results to the user, who then selects a voice actor from the displayed list.
[1842] Step 3:
[1843] The user (video creator) accesses the details page of the voice actor they selected and enters a prompt sentence (e.g., "Buy now and get a special discount!"). This entered text is sent to the server. At the same time, the emotion recognition engine (Emotion AI) collects emotional data from the user's facial expressions and voice in real time and sends it to the server.
[1844] Step 4:
[1845] The server sends prompts to a text-to-speech engine based on the acquired emotional data and the input text. The text-to-speech engine then generates appropriate, emotionally rich speech data. In this process, the emotional data sent from the emotion recognition engine is reflected in the speech generation.
[1846] Step 5:
[1847] The server converts the generated audio data into an appropriate format and provides it to the user's (video creator's) device. The user then downloads the generated audio data and integrates it into their own video content or advertisements. This results in the creation of a high-quality video work that includes richly expressive audio.
[1848] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1849] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1850] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1851] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1852] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1853] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1854] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1855] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1856] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1857] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1858] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1859] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1860] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1861] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1862] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1863] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1864] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1865] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1866] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1867] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1868] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1869] The following is further disclosed regarding the above embodiment.
[1870] (Claim 1)
[1871] A system for video creators to generate audio data, comprising:
[1872] A means for voice actors to register their own voice data and terms of use;
[1873] A way for filmmakers to search for voice actors using keywords;
[1874] means for converting input text into speech based on voice data of a selected voice actor;
[1875] means for incorporating the generated audio data into the video;
[1876] A system including:
[1877] (Claim 2)
[1878] 10. The system of claim 1, further comprising means for storing voice data of voice actors and terms of use in a database.
[1879] (Claim 3)
[1880] 10. The system of claim 1, wherein the selection of voice actors and generation of voice data is performed using voice generation AI.
[1881] "Example 1"
[1882] (Claim 1)
[1883] A system for video creators to generate audio data, comprising:
[1884] A means for voice actors to register their own voice data and terms of use;
[1885] A way for filmmakers to search for voice actors using keywords,
[1886] means for converting input text into speech based on voice data of a selected voice actor;
[1887] a means for incorporating the generated audio data into a video work;
[1888] a means for converting input text into speech using a speech generation AI model;
[1889] means for providing the generated voice data to a user;
[1890] A system including:
[1891] (Claim 2)
[1892] 10. The system of claim 1, further comprising means for storing voice data of voice actors and terms of use in a database.
[1893] (Claim 3)
[1894] 10. The system of claim 1, wherein the selection of voice actors and generation of voice data is performed using voice generation AI.
[1895] "Application Example 1"
[1896] (Claim 1)
[1897] 1. A content distribution system for generating audio data, comprising:
[1898] A means for a voice provider to register his / her voice data and terms of use;
[1899] a means for content creators to search for audio providers using keywords;
[1900] means for converting input text into speech based on speech data of a selected speech provider;
[1901] means for incorporating the generated audio data into content in real time;
[1902] means for providing an interactive experience through a visual display device;
[1903] A system including:
[1904] (Claim 2)
[1905] The system further comprises means for storing the voice data and the terms of use of the voice provider in a database;
[1906] 10. The system of claim 1.
[1907] (Claim 3)
[1908] The selection of voice providers and the generation of voice data are performed using a generative AI model;
[1909] 10. The system of claim 1.
[1910] "Example 2: Combining Emotion Engines"
[1911] (Claim 1)
[1912] A means for voice actors to register their own voice data and terms of use;
[1913] A way for filmmakers to search for voice actors using keywords;
[1914] means for converting input text into speech based on voice data of a selected voice actor;
[1915] means for recognizing a user's emotion and generating voice data based on the emotion;
[1916] means for incorporating the generated audio data into the video;
[1917] A system including:
[1918] (Claim 2)
[1919] 10. The system of claim 1, further comprising means for storing voice data of voice actors and terms of use in a database.
[1920] (Claim 3)
[1921] 10. The system of claim 1, wherein the selection of voice actors and the generation of voice data are performed using voice generation artificial intelligence.
[1922] "Application example 2 when combining emotion engines"
[1923] (Claim 1)
[1924] A system for video creators to generate audio data, comprising:
[1925] A means for voice actors to register their own voice data and terms of use;
[1926] A way for filmmakers to search for voice actors using keywords;
[1927] means for converting input text into speech based on voice data of a selected voice actor;
[1928] means for recognizing a user's emotion and generating voice data based on the emotion;
[1929] means for incorporating the generated audio data into the video;
[1930] A system including:
[1931] (Claim 2)
[1932] 10. The system of claim 1, further comprising means for storing voice data of voice actors and terms of use in a database.
[1933] (Claim 3)
[1934] 10. The system of claim 1, wherein the voice actor is selected and the voice data is generated using a voice generation AI, and the generated voice data reflects the user's emotions. [Explanation of symbols]
[1935] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A system for video creators to generate audio data, comprising: A means for voice actors to register their own voice data and terms of use; A way for filmmakers to search for voice actors using keywords; means for converting input text into speech based on voice data of a selected voice actor; means for incorporating the generated audio data into the video; A system including:
2. 10. The system of claim 1, further comprising means for storing voice data of voice actors and terms of use in a database.
3. 10. The system of claim 1, wherein the selection of voice actors and the generation of voice data are performed using a voice generation AI.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A