system

A licensing-based system for audio generation models addresses the loss of income for voice providers and ensures reliable audio by attaching official authentication, allowing voice providers to earn fair compensation and users to access authentic audio.

JP2026068486APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

The spread of generative AI technology for voice mimicry has led to a loss of income sources for voice providers and a lack of reliable mechanisms to determine the authenticity of generated voices, necessitating a system for providing official and reliable audio data.

Method used

A system that enters into licensing agreements with audio providers, constructs audio generation models using their data, and ensures fair compensation through usage fees, attaching official authentication to generated audio.

Benefits of technology

Enables voice providers to earn a sustainable income while ensuring users receive high-quality, reliable audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068486000001_ABST
    Figure 2026068486000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of concluding a license agreement with an audio provider to permit the provision of audio data, A means for creating a speech generation model based on the voice data of a voice provider, A means for generating voice data using the voice generation model in response to a request from a user terminal, A means of attaching official authentication information to the generated audio data, A method for collecting usage fees from users and distributing them to voice providers as license fees, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] With the development of generative AI technology, an environment is being created where it is easy to generate voice data that mimics the voices of famous voice providers. However, the spread of this technology has led to the problem of the loss of the income sources that voice providers should originally obtain. Also, since there is currently a lack of a reliable mechanism for determining whether the generated voice is official, there is a demand for the provision of official voice data that consumers can use with confidence.

Means for Solving the Problems

[0005] This invention provides a system for generating official audio data by entering into a license agreement with an audio provider and constructing an audio generation model based on that audio data. This system generates audio data using the audio generation model in response to a request from a user terminal and assigns official authentication information to the generated audio. Furthermore, by collecting usage fees from users and distributing them to audio providers as license fees, the system enables audio providers to receive fair compensation. This allows for the distribution of official and reliable audio data, preventing potential revenue losses for audio providers.

[0006] "Audio data" refers to data that records the voice of the voice provider in digital format.

[0007] A "license agreement" is a contract entered into by an audio provider to grant permission for the use of audio data.

[0008] A "voice provider" is an individual or organization that provides their voice as a model for a voice-generating AI.

[0009] A "speech generation model" is an algorithm or program that generates speech using a computer based on speech data.

[0010] A "user terminal" is a device that can connect to a voice generation system and send requests.

[0011] "Official authentication information" refers to data added to indicate that the generated audio data is official and legitimate.

[0012] "Usage fees" refer to the fees that users pay for using audio data.

[0013] A "license fee" is the compensation that an audio provider receives for granting permission to use their audio data. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a system that allows voice providers to profit from the use of voice generation models based on their own voice data. This system enters into licensing agreements with voice providers, collects their voice data, and develops voice generation models. The developed models are integrated into user applications and accept voice generation requests from user terminals.

[0036] The server generates audio data using a speech generation model based on requests sent from the user's terminal. During this process, official authentication information is attached to the generated audio data, enabling the provision of reliable audio to the user. The user receives the generated audio through the application and pays a usage fee based on its use. The paid usage fees are collected by the server and distributed to the audio providers as license fees.

[0037] As a concrete example, if a user wants to generate a message in the voice of a specific character, the user launches the application and enters the desired text. The device sends this request to the server, which uses a speech generation model to create the voice. The generated voice is immediately sent to the user's device, and its quality and reliability are guaranteed by official authentication information. After the user confirms the voice, the usage fee is processed within the application, and the license fee is returned to the voice provider, who is the voice actor. This system allows voice actors to earn a sustainable income by utilizing their voices, and fans can enjoy messages using the voice actors' voices with peace of mind.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] The server will enter into a license agreement with the audio provider and define the procedure for the audio provider to provide their audio data to the server.

[0041] Step 2:

[0042] The server applies a speech processing algorithm to the received audio data in order to begin developing a speech generation model.

[0043] Step 3:

[0044] The server integrates the developed speech generation model into an application accessible to users and distributes the application.

[0045] Step 4:

[0046] Users download the application and create and configure a user account when they begin using it.

[0047] Step 5:

[0048] The user enters the text they want to have spoken in the application's interface and initiates the request.

[0049] Step 6:

[0050] The terminal sends the entered text information to the server and requests the generation of audio data.

[0051] Step 7:

[0052] The server generates speech data based on the input text using a speech generation model.

[0053] Step 8:

[0054] The server assigns official authentication information to the generated audio data, guaranteeing the reliability of the audio data.

[0055] Step 9:

[0056] The server sends voice data containing official authentication information to the terminal, making it available to the user.

[0057] Step 10:

[0058] The user reviews the received audio data and then proceeds to pay the usage fee for its use.

[0059] Step 11:

[0060] The device sends payment information for the usage fee to the server, and the payment is confirmed.

[0061] Step 12:

[0062] The server collects usage fees and distributes a portion of them to the voice providers as licensing fees, thereby returning profits to them.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Conventional voice generation systems have problems with adequately protecting the rights of voice providers and with users being unable to obtain reliable audio information. Furthermore, there is a lack of mechanisms to guarantee the reliability and quality of the voice generation models used, making it difficult to provide reliable value to users. In addition, mechanisms for efficiently managing and fairly distributing royalties to rights holders when users utilize specific voices are immature.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for concluding a rights agreement with the speaker to permit the provision of acoustic information, means for creating a speech generation algorithm based on the speaker's acoustic information, and means for generating acoustic information using the speech generation algorithm in response to a request from a user device. This ensures the reliability and quality of the acoustic information through official authentication information, enabling reliable speech generation for users and protecting the rights of the speaker, as well as effective management and distribution of royalties.

[0068] "Acoustic information" refers to data that digitally represents the characteristics of sound propagating through space.

[0069] A "rights agreement" refers to a legal agreement in which an audio provider grants permission for a third party to use their audio information.

[0070] A "speaker" refers to someone who provides their own sonic information.

[0071] A "speech generation algorithm" refers to a computational method or process for generating new acoustic information based on existing acoustic information.

[0072] A "user device" refers to an electronic device that a user can directly operate and that has the function of generating, receiving, and playing back acoustic information.

[0073] A "request" refers to a command sent from a user device to a server requesting specific processing or information provision.

[0074] "Official authentication information" refers to the digital signatures and metadata attached to prove that the generated audio information is trustworthy.

[0075] "Presentation text" refers to the text data input into a speech generation algorithm, and is a sentence that indicates the content that forms the basis of the generated acoustic information.

[0076] An "input / output mechanism" refers to an interface that allows a user to input or output data through their device.

[0077] "Transmission" refers to the process of sending generated acoustic information to the user's device.

[0078] This invention is realized by a speech generation system based on the acoustic information of voice providers. The server enters into a rights agreement with the voice providers and collects their acoustic information. The collected acoustic information is used to develop a speech generation algorithm. This algorithm is trained using machine learning frameworks such as TENSORFLOW® and PyTorch, and is built as a model capable of generating new acoustic information.

[0079] If a user wants to generate audio information, they use an application installed on their device. The user enters the desired text into this application. Specifically, they might enter a prompt such as, "Convert the following text into speech in the voice of the specified character: 'Good morning! What fun things await me today?'" This entered data is then sent from the device to the server.

[0080] The server inputs the received text data into a configured speech generation algorithm and uses a generation AI model to generate audio information. During this process, the server guarantees the reliability and quality of the information by attaching official authentication information to the generated audio. Finally, the generated audio information is sent from the server to the user's terminal, where the user can review and use the audio.

[0081] Through this invention, audio providers can earn revenue while protecting their rights, and users can easily access high-quality, reliable audio information.

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The server enters into a rights agreement with the voice provider and collects acoustic information. The input for this step is the acoustic information provided by the voice provider. The server stores the acoustic information as digital data in preparation for later model development.

[0085] Step 2:

[0086] The server uses the collected acoustic information to train a speech generation algorithm. The input for this step is the stored acoustic information as digital data. The server leverages a machine learning platform (e.g., TensorFlow) to train a generative AI model. The output is a generative AI model equipped with the speech generation algorithm.

[0087] Step 3:

[0088] The user enters the desired text into the application on their device. This input is text data based on the user's preferences. The device uses this data as a prompt and sends it to the server as a speech generation request.

[0089] Step 4:

[0090] The server provides the received text data to the AI ​​model as a prompt. The input for this step is the text data sent by the user. The server uses a speech generation algorithm to generate new acoustic information based on this text. The output is the generated acoustic information.

[0091] Step 5:

[0092] The server adds official authentication information to the generated acoustic information. This step adds metadata indicating the reliability of the acoustic information. The input to this process is the acoustic information resulting from speech generation, and the output is the acoustic information with official authentication information added.

[0093] Step 6:

[0094] The server sends audio information with official authentication credentials to the user's device. The input in this step is authenticated audio information. The server uses real-time communication technology to stream the data, enabling audio playback on the device.

[0095] Step 7:

[0096] The user confirms the audio information received on the device. After confirmation, the user pays the usage fee. The input is the played audio information and the user's decision to use it. The device sends the payment information to the server, and the output is the successfully processed payment data.

[0097] Step 8:

[0098] The server distributes the collected royalties to the voice providers. The input for this step is payment data from the user. Based on the contract with the voice provider, the server calculates the royalties, and the output is the distribution of appropriate license fees to the voice provider.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] The challenge is to provide a system that delivers reliable audio content to audio content enthusiasts and enables audio providers to continuously generate revenue. Another challenge is to enable users to instantly generate and obtain audio information tailored to their preferences.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data; means for creating a voice generation model based on the voice provider's voice data; means for generating acoustic information using the voice generation model in response to a request from a user terminal; means for attaching official authentication information to the generated acoustic information; means for collecting usage fees from users and distributing them to voice providers as license fees; means for providing an information input area for generating acoustic information; and means for transmitting the generated acoustic information to an information terminal. As a result, users can obtain reliable acoustic content with peace of mind, and voice providers can achieve stable monetization.

[0104] A "voice provider" is an individual or organization that provides their own voice data and cooperates in the development of a voice generation model.

[0105] A "license agreement" is a legal agreement entered into by an audio provider to permit a third party to use the audio data.

[0106] A "speech generation model" is an algorithm or program built on speech data from a speech provider to generate speech corresponding to specified text.

[0107] A "user terminal" is an electronic device owned by a user and used to generate and receive acoustic information.

[0108] "Official authentication information" refers to digital information indicating reliability, which is assigned to guarantee that the generated audio information is genuine.

[0109] A "usage fee" is money paid by a user as compensation for acquiring and using audio information.

[0110] "Acoustic information" refers to audio data created using a speech generation model.

[0111] The "information input area" is the interface section provided for users to input the text and parameters necessary for generating acoustic information.

[0112] An "information terminal" is a digital device used to receive and play back generated audio information.

[0113] This invention is a system that develops a speech generation model using speech data from a speech provider and provides reliable acoustic information to the user. The server comprises several key components for this purpose. It creates a speech generation model based on speech data obtained from the speech provider and operates efficiently using cloud-based computing resources. This model is built using a deep learning framework such as TensorFlow. The server also leverages Google Cloud's Text to Speech API to attach official authentication information to the generated acoustic information.

[0114] The user terminal refers to an information device such as a smartphone or tablet, which sends a voice generation request through an application. The user can enter text into an information input area within the application and select their preferred voice provider. This request is sent to a server via the internet, where the necessary audio information is generated by a designated generation AI model. The generated audio information is returned to the user terminal, where its quality is verified by official authentication information. Users can then enjoy the audio content with peace of mind.

[0115] A concrete example of this use case is when a user creates a message to a friend using the voice of their favorite voice actor. The user enters a text message such as "Hello, I'm the AI ​​assistant. Have a great day!" into the application's text input area, and the application generates a message in the voice of that voice actor. Based on this prompt, the server responds immediately and provides reliable audio information to the user's device.

[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0117] Step 1:

[0118] The user launches the application on their smartphone or tablet and enters the text of the voice message they want to generate into the information input area. This entered text is the input data used in the next step of the process.

[0119] Step 2:

[0120] The terminal sends the voice provider and text information to the server according to the user's selection. This information is processed on the server side as basic data for selecting a speech generation model and for text-to-speech conversion.

[0121] Step 3:

[0122] The server selects the optimal speech generation model based on the received text and information from the voice provider. This model operates using the TensorFlow framework and generates acoustic information in the specified voice using a deep learning algorithm.

[0123] Step 4:

[0124] The selected speech generation model uses Google Cloud's Text to Speech API to convert text into speech data. This process involves parsing the text data and generating the corresponding speech waveform. The converted speech data is the output.

[0125] Step 5:

[0126] The server assigns official authentication information to the generated audio data. This information is generated using digital signatures and other methods to guarantee the authenticity of the audio.

[0127] Step 6:

[0128] The server sends authenticated audio data to the user's device. This data is then played back by the user to check its quality and content.

[0129] Step 7:

[0130] The user receives the audio data and plays it within the application. The user checks the sound and verifies that the audio is correctly generated in the voice selected by the voice provider.

[0131] Step 8:

[0132] The terminal calculates usage fees based on the use of voice data and sends that information back to the server. The server then collects the fees based on this information and distributes them to the voice providers as license fees.

[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0134] This invention enables the generation of voices that respond to the user's emotions by integrating an emotion engine into a system that provides voices generated using voice data from voice providers to the user's terminal. The system enters into a license agreement with a voice provider and constructs a voice generation model based on that voice data.

[0135] The server receives text requests and emotion data sent from the user's terminal and first analyzes the user's emotional state using an emotion engine. The emotion engine evaluates the user's emotions in real time using user input and data obtained from the terminal's sensors. This emotion data is then input into a speech generation model, which adjusts the speech according to the user's emotions.

[0136] For example, if a user wants to generate character dialogue based on feelings of joy through an application, the text entered by the user is sent to the server. The server uses an emotion engine to evaluate the emotion data obtained from the device and incorporates the results into the voice generation process, thereby generating voice with a nuance of joy. This voice is then accompanied by official authentication information and sent to the user's device.

[0137] Users review the received audio and pay a usage fee within the application based on their usage. After the server collects the usage fee, it returns it to the audio provider as a license fee. This system allows audio providers to earn revenue based on their audio assets, and users can confidently use audio content that resonates with their emotions.

[0138] The following describes the processing flow.

[0139] Step 1:

[0140] The user launches the application and accesses the interface for speech generation. The user enters the content of the speech they want to generate as text.

[0141] Step 2:

[0142] The device acquires text data entered by the user and also collects the user's emotional data through emotion recognition sensors (camera, microphone, biometric data, etc.).

[0143] Step 3:

[0144] The device sends the collected text and emotion data to the server and requests speech generation.

[0145] Step 4:

[0146] The server uses an emotion engine to analyze the received emotion data and evaluate the user's specific emotions (joy, sadness, anger, etc.).

[0147] Step 5:

[0148] The server adjusts the parameters of the speech generation model based on the user's emotion evaluation results. This ensures that the emotion is reflected in the tone and intonation of the generated speech.

[0149] Step 6:

[0150] The server converts text data into speech data using a finely tuned speech generation model.

[0151] Step 7:

[0152] The server assigns official authentication information to the generated audio data, creating reliable and rights-protected audio data.

[0153] Step 8:

[0154] The server sends audio data containing official authentication information to the terminal and provides it to the user.

[0155] Step 9:

[0156] The user reviews the received audio data. If the user is satisfied with the content and emotional expression of the audio, they decide to pay a usage fee through the application.

[0157] Step 10:

[0158] The device sends payment information to the server and processes the usage fee.

[0159] Step 11:

[0160] The server collects usage fees from users and returns them to the voice providers as license fees in proportion to the licensing agreement.

[0161] (Example 2)

[0162] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0163] Conventional voice generation systems have difficulty generating voices that respond to user emotions, and have failed to accurately reflect the emotional nuances desired by the user. Furthermore, they lacked mechanisms for effectively utilizing voice providers' voice data and managing licenses appropriately. It is necessary to solve these problems and achieve more natural and emotionally resonant voice generation.

[0164] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0165] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data, means for analyzing the user's emotional state based on input data from a terminal, and means for generating voice data using a voice generation model in response to requests from the user terminal and the analyzed emotional state. This enables voice generation that responds to the user's emotions, and allows for the provision of voice content that meets the diverse needs of users while appropriately managing the voice provider's voice assets.

[0166] A "license agreement for providing audio data" is a legal agreement in which an audio provider grants others the right to use their audio data and receives compensation for it.

[0167] A "speech generation model" is a computational algorithm or mathematical model that takes audio data as input and generates new audio based on that data.

[0168] "Means for analyzing a user's emotional state based on input data from a device" refers to systems or methods for identifying a user's current emotions based on data (voice, video, text, etc.) obtained from the device the user is using.

[0169] A "user terminal request" refers to instructions or information sent by a user through their device to request a specific service or function.

[0170] "Means of providing official authentication information" refers to the process of adding information to generated data or content to prove its legitimacy and origin.

[0171] "Means of distributing licensing fees to audio providers" refers to methods or systems for appropriately distributing usage fees received from users to providers who hold the rights to the audio data.

[0172] This invention is a system that utilizes voice data from voice providers to generate voices that respond to the user's emotions. Specifically, a license agreement is concluded with the voice provider, and a voice generation model is built based on that voice data. The server receives text requests and emotion data from the user terminal. The emotion data is obtained from user input and terminal sensors and analyzed by the emotion engine. Based on the results of this analysis, the voice generation model generates voices that match the user's emotions.

[0173] The system implementation utilizes various hardware and software. Specifically, the server runs on a cloud platform, and speech generation uses generative AI technology and a speech synthesis API. For emotion analysis, an emotion analysis API is used to evaluate the user's emotions in real time. The user terminal inputs a voice request and runs an application to receive and play the generated voice. This allows users to easily obtain voice content that matches their desired emotion.

[0174] For example, if a user wants to generate an audio message that is "cheerful and encouraging," they would enter the following prompt.

[0175] Example of a prompt: "Let's do our best today. Please speak in a cheerful and bright voice."

[0176] Based on this input, the server performs appropriate sentiment analysis and uses a speech generation model to generate speech that reflects an "energetic tone," which is then delivered to the user's terminal. This system allows users to quickly obtain speech that matches their emotions, and enables speech providers to secure revenue by utilizing their speech assets.

[0177] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0178] Step 1:

[0179] The user uses their device to input prompt text for voice generation. For example, they might enter text such as "Please create an uplifting message" into the application. The entered text is sent from the device to the server. This data is used to verify whether the entered text aligns with the user's expected emotional response.

[0180] Step 2:

[0181] The server receives emotional data from the terminal's sensors (e.g., a facial recognition camera or voice tone microphone) along with the prompt message received from the terminal. This allows the server to understand the user's current emotional state and use it as data necessary for analysis. The server uses an emotion engine to analyze the acquired data and evaluate the user's emotional state. This evaluation result becomes the output used to adjust the speech generation model.

[0182] Step 3:

[0183] The server inputs the evaluated emotion data and prompt text into a speech generation model. This model utilizes generative AI technology to process the user's emotions and text to generate appropriate speech. The speech generation model outputs speech data that reflects the emotions, incorporating the nuances the user expects into the speech.

[0184] Step 4:

[0185] The server adds official authentication information to the generated audio data, guaranteeing that the original audio data is legitimate. This audio data is then ready to be sent to the user's terminal. The addition of official authentication information is a specific action taken to ensure the reliability of the data.

[0186] Step 5:

[0187] The user's device receives audio data sent from the server, and the user plays the audio using an application. The user can review the audio and request regeneration if necessary. This process allows the user to verify the audio quality and emotional appropriateness.

[0188] Step 6:

[0189] Users pay a fee for the use of the generated audio, and the server manages this payment information. Once the payment processing is complete, the server distributes the license fee to the audio provider. This allows the audio provider to earn revenue through their audio data. The payment and distribution stages are crucial processes for ensuring the overall profitability of the service and the convenience of the user.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0192] In audio content distribution, there is a growing need to provide users with a more immersive experience by offering audio that responds to their emotional state. A method for efficiently delivering services that meet this requirement is necessary. Furthermore, fair compensation for audio providers must also be considered.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes means for concluding a licensing agreement with a voice provider to provide voice data, means for creating a voice generation model based on the voice data, and means for analyzing the user's emotional state and acquiring emotional data. This makes it possible to generate and provide voice data that corresponds to the user's emotions.

[0195] A "voice provider" is a person who has the right to provide voice data and whose data is used to construct a voice generation model.

[0196] A "licensing agreement" is a contract that forms an agreement with the audio provider for the use of specific rights.

[0197] A "speech generation model" is an algorithm or program that generates speech that corresponds to the user's emotions based on the speech data of a speech provider.

[0198] An "emotion analysis device" is a device that analyzes a user's emotional state based on information provided by the user and data from terminal sensors.

[0199] "Voice data generation" is the process of generating voice data to be provided to the user using a voice generation model based on acquired emotional data.

[0200] "Official authentication information" refers to identification information assigned to prove that the generated audio data is legitimate.

[0201] "Usage fees" refer to the fees that users must pay to use the generated audio data.

[0202] "License fees" refer to the payment made to the voice provider for the right to use the voice data and build the model.

[0203] To implement this invention, a voice provider provides voice data to a server, and the server constructs a voice generation model based on that data. A licensing agreement is concluded between the voice provider and the server, securing the right to use the voice data. The server receives emotion data transmitted from the user terminal via an emotion analysis device. This emotion data is acquired through user input and sensors and is used to analyze the user's emotional state in real time.

[0204] The server analyzes the user's emotional state and uses a speech generation model to generate audio data corresponding to that emotional data. The generated audio data is accompanied by official authentication information to guarantee its authenticity. The speech generation model is developed using Python and related libraries (PyTorch, TensorFlow). The generated audio data is transmitted to the user's terminal via the internet.

[0205] The user confirms that the audio data has been provided on their device and pays the usage fee through a digital payment system. The server distributes the usage fee to the audio provider as a licensing fee. This allows users to use audio content that fits their emotions, and audio providers to receive compensation for the audio data they provide.

[0206] For example, when a user is listening to a horror novel in an audiobook application, if the emotion analyzer recognizes the user's level of tension, the voice generated by the model will be adjusted to a gentler tone. This mechanism improves the accuracy of emotion-based voice generation and enhances the user experience. The voice generation process can be controlled using a prompt such as, "Generate voice reading the next text in a manner appropriate for when the user's stress level is high."

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The user terminal requests audio content and collects data related to the user's emotions during the process. It receives user voice input and sensor information as input, processes it with an emotion analysis device, and outputs emotion data. Specifically, it uses the user terminal's microphone and camera to capture the user's facial expressions and voice tone in real time.

[0210] Step 2:

[0211] The server receives emotion data sent from the user's terminal. Based on the input emotion data, it uses an emotion analysis engine to analyze the user's emotional state and outputs the results in a data format. Specifically, it analyzes the obtained emotion information and evaluates the user's emotional state (e.g., joy, sadness, tension).

[0212] Step 3:

[0213] The server constructs a speech generation model based on the analyzed emotion data and the user's request text. The emotion data and text are input to the speech generation model, and the adjusted speech data is output. Specifically, the generation AI model built on the server generates a speech waveform corresponding to the emotion.

[0214] Step 4:

[0215] The server assigns official authentication information to the generated audio data. It labels the audio data with authentication information in text format and outputs the result. Specifically, it adds a digital signature to the generated audio file to verify its authenticity.

[0216] Step 5:

[0217] The server sends audio data with official authentication information attached to the user's terminal. The audio data is transferred using a communication protocol, and the audio file is output to the user's terminal. Specifically, the audio data is streamed in real time over the internet.

[0218] Step 6:

[0219] The user receives and confirms the voice data on their device and pays the usage fee. They verify that the voice is as expected as input and pay the fee as output through a digital payment system. Specifically, they confirm the billing information on the application and complete the payment process.

[0220] Step 7:

[0221] The server distributes the collected royalties to the voice providers as licensing fees. Based on the royalty data received as input, it generates payment orders to the voice providers as output. Specifically, it uses an electronic payment platform to automatically process transfers to the voice providers.

[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0225] [Second Embodiment]

[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0238] This invention provides a system that allows voice providers to profit from the use of voice generation models based on their own voice data. This system enters into licensing agreements with voice providers, collects their voice data, and develops voice generation models. The developed models are integrated into user applications and accept voice generation requests from user terminals.

[0239] The server generates audio data using a speech generation model based on requests sent from the user's terminal. During this process, official authentication information is attached to the generated audio data, enabling the provision of reliable audio to the user. The user receives the generated audio through the application and pays a usage fee based on its use. The paid usage fees are collected by the server and distributed to the audio providers as license fees.

[0240] As a concrete example, if a user wants to generate a message in the voice of a specific character, the user launches the application and enters the desired text. The device sends this request to the server, which uses a speech generation model to create the voice. The generated voice is immediately sent to the user's device, and its quality and reliability are guaranteed by official authentication information. After the user confirms the voice, the usage fee is processed within the application, and the license fee is returned to the voice provider, who is the voice actor. This system allows voice actors to earn a sustainable income by utilizing their voices, and fans can enjoy messages using the voice actors' voices with peace of mind.

[0241] The following describes the processing flow.

[0242] Step 1:

[0243] The server will enter into a license agreement with the audio provider and define the procedure for the audio provider to provide their audio data to the server.

[0244] Step 2:

[0245] The server applies a speech processing algorithm to the received audio data in order to begin developing a speech generation model.

[0246] Step 3:

[0247] The server integrates the developed speech generation model into an application accessible to users and distributes the application.

[0248] Step 4:

[0249] Users download the application and create and configure a user account when they begin using it.

[0250] Step 5:

[0251] The user enters the text they want to have spoken in the application's interface and initiates the request.

[0252] Step 6:

[0253] The terminal sends the entered text information to the server and requests the generation of audio data.

[0254] Step 7:

[0255] The server generates speech data based on the input text using a speech generation model.

[0256] Step 8:

[0257] The server assigns official authentication information to the generated audio data, guaranteeing the reliability of the audio data.

[0258] Step 9:

[0259] The server sends voice data containing official authentication information to the terminal, making it available to the user.

[0260] Step 10:

[0261] The user reviews the received audio data and then proceeds to pay the usage fee for its use.

[0262] Step 11:

[0263] The device sends payment information for the usage fee to the server, and the payment is confirmed.

[0264] Step 12:

[0265] The server collects usage fees and distributes a portion of them to the voice providers as licensing fees, thereby returning profits to them.

[0266] (Example 1)

[0267] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0268] Conventional voice generation systems have problems with adequately protecting the rights of voice providers and with users being unable to obtain reliable audio information. Furthermore, there is a lack of mechanisms to guarantee the reliability and quality of the voice generation models used, making it difficult to provide reliable value to users. In addition, mechanisms for efficiently managing and fairly distributing royalties to rights holders when users utilize specific voices are immature.

[0269] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0270] In this invention, the server includes means for concluding a rights agreement with the speaker to permit the provision of acoustic information, means for creating a speech generation algorithm based on the speaker's acoustic information, and means for generating acoustic information using the speech generation algorithm in response to a request from a user device. This ensures the reliability and quality of the acoustic information through official authentication information, enabling reliable speech generation for users and protecting the rights of the speaker, as well as effective management and distribution of royalties.

[0271] "Acoustic information" refers to data that digitally represents the characteristics of sound propagating through space.

[0272] A "rights agreement" refers to a legal agreement in which an audio provider grants permission for a third party to use their audio information.

[0273] A "speaker" refers to someone who provides their own sonic information.

[0274] A "speech generation algorithm" refers to a computational method or process for generating new acoustic information based on existing acoustic information.

[0275] A "user device" refers to an electronic device that a user can directly operate and that has the function of generating, receiving, and playing back acoustic information.

[0276] A "request" refers to a command sent from a user device to a server requesting specific processing or information provision.

[0277] "Official authentication information" refers to the digital signatures and metadata attached to prove that the generated audio information is trustworthy.

[0278] "Presentation text" refers to the text data input into a speech generation algorithm, and is a sentence that indicates the content that forms the basis of the generated acoustic information.

[0279] An "input / output mechanism" refers to an interface that allows a user to input or output data through their device.

[0280] "Transmission" refers to the process of sending generated acoustic information to the user's device.

[0281] This invention is realized by a speech generation system based on the acoustic information of voice providers. The server enters into a rights agreement with the voice providers and collects their acoustic information. The collected acoustic information is used to develop a speech generation algorithm. This algorithm is trained using machine learning frameworks such as TensorFlow and PyTorch and built as a model capable of generating new acoustic information.

[0282] When a user wants to generate acoustic information, the application installed on the terminal is used. The user inputs the desired text in this application. In particular, as a prompt sentence, "Please convert the following text into voice in the voice of the specified character: 'Good morning! What fun things are waiting for us today?'" etc. is input. The input data is sent from the terminal to the server.

[0283] The server inputs the received text data into the set voice generation algorithm and utilizes the generation AI model to generate acoustic information. In this process, the server guarantees the reliability and quality of the information by attaching official authentication information to the generated acoustic information. Finally, the generated acoustic information is sent from the server to the user's terminal, and the user can confirm and use the voice.

[0284] Through this invention, the voice provider can obtain benefits while protecting its own rights, and the user can easily utilize high-quality and reliable acoustic information.

[0285] The flow of the specific process in Example 1 will be described with reference to FIG. 11.

[0286] Step 1:

[0287] The server concludes a rights contract with the voice provider and collects acoustic information. The input of this step is the acoustic information provided by the voice provider. The server stores the acoustic information as digital data in preparation for subsequent model development.

[0288] Step 2:

[0289] The server uses the collected acoustic information to train a speech generation algorithm. The input for this step is the stored acoustic information as digital data. The server leverages a machine learning platform (e.g., TensorFlow) to train a generative AI model. The output is a generative AI model equipped with the speech generation algorithm.

[0290] Step 3:

[0291] The user enters the desired text into the application on their device. This input is text data based on the user's preferences. The device uses this data as a prompt and sends it to the server as a speech generation request.

[0292] Step 4:

[0293] The server provides the received text data to the AI ​​model as a prompt. The input for this step is the text data sent by the user. The server uses a speech generation algorithm to generate new acoustic information based on this text. The output is the generated acoustic information.

[0294] Step 5:

[0295] The server adds official authentication information to the generated acoustic information. This step adds metadata indicating the reliability of the acoustic information. The input to this process is the acoustic information resulting from speech generation, and the output is the acoustic information with official authentication information added.

[0296] Step 6:

[0297] The server sends audio information with official authentication credentials to the user's device. The input in this step is authenticated audio information. The server uses real-time communication technology to stream the data, enabling audio playback on the device.

[0298] Step 7:

[0299] The user confirms the audio information received on the device. After confirmation, the user pays the usage fee. The input is the played audio information and the user's decision to use it. The device sends the payment information to the server, and the output is the successfully processed payment data.

[0300] Step 8:

[0301] The server distributes the collected royalties to the voice providers. The input for this step is payment data from the user. Based on the contract with the voice provider, the server calculates the royalties, and the output is the distribution of appropriate license fees to the voice provider.

[0302] (Application Example 1)

[0303] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0304] The challenge is to provide a system that delivers reliable audio content to audio content enthusiasts and enables audio providers to continuously generate revenue. Another challenge is to enable users to instantly generate and obtain audio information tailored to their preferences.

[0305] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0306] In this invention, the server includes means for concluding a license contract with a voice provider to permit the provision of voice data, means for creating a voice generation model based on the voice data of the voice provider, means for generating acoustic information using the voice generation model in response to a request from a user terminal, means for attaching official authentication information to the generated acoustic information, means for collecting usage fees from users and distributing them as license fees to the voice provider, means for providing an information input area for generating acoustic information, and means for transmitting the generated acoustic information to an information terminal. As a result, users can obtain reliable acoustic content with confidence, and stable monetization by voice providers becomes possible.

[0307] A "voice provider" is an individual or organization that provides its own voice data and cooperates in the development of a voice generation model.

[0308] A "license contract" is a legal agreement concluded by a voice provider to permit a third party to use voice data.

[0309] A "voice generation model" is an algorithm or program constructed based on the voice data of a voice provider and that generates voice corresponding to specified text.

[0310] A "user terminal" is an electronic device owned by a user and used for generating and receiving acoustic information.

[0311] "Official authentication information" is digital information indicating reliability that is attached to guarantee that the generated acoustic information is legitimate.

[0312] A "usage fee" is money paid by a user as consideration for obtaining and using acoustic information.

[0313] "Acoustic information" is voice data created using a voice generation model.

[0314] The "information input area" is the interface section provided for users to input the text and parameters necessary for generating acoustic information.

[0315] An "information terminal" is a digital device used to receive and play back generated audio information.

[0316] This invention is a system that develops a speech generation model using speech data from a speech provider and provides reliable acoustic information to the user. The server comprises several key components for this purpose. It creates a speech generation model based on speech data obtained from the speech provider and operates efficiently using cloud-based computing resources. This model is built using a deep learning framework such as TensorFlow. The server also leverages Google Cloud's Text to Speech API to attach official authentication information to the generated acoustic information.

[0317] The user terminal refers to an information device such as a smartphone or tablet, which sends a voice generation request through an application. The user can enter text into an information input area within the application and select their preferred voice provider. This request is sent to a server via the internet, where the necessary audio information is generated by a designated generation AI model. The generated audio information is returned to the user terminal, where its quality is verified by official authentication information. Users can then enjoy the audio content with peace of mind.

[0318] A concrete example of this use case is when a user creates a message to a friend using the voice of their favorite voice actor. The user enters a text message such as "Hello, I'm the AI ​​assistant. Have a great day!" into the application's text input area, and the application generates a message in the voice of that voice actor. Based on this prompt, the server responds immediately and provides reliable audio information to the user's device.

[0319] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0320] Step 1:

[0321] The user launches the application on their smartphone or tablet and enters the text of the voice message they want to generate into the information input area. This entered text is the input data used in the next step of the process.

[0322] Step 2:

[0323] The terminal sends the voice provider and text information to the server according to the user's selection. This information is processed on the server side as basic data for selecting a speech generation model and for text-to-speech conversion.

[0324] Step 3:

[0325] The server selects the optimal speech generation model based on the received text and information from the voice provider. This model operates using the TensorFlow framework and generates acoustic information in the specified voice using a deep learning algorithm.

[0326] Step 4:

[0327] The selected speech generation model uses Google Cloud's Text to Speech API to convert text into speech data. This process involves parsing the text data and generating the corresponding speech waveform. The converted speech data is the output.

[0328] Step 5:

[0329] The server assigns official authentication information to the generated audio data. This information is generated using digital signatures and other methods to guarantee the authenticity of the audio.

[0330] Step 6:

[0331] The server sends authenticated audio data to the user's device. This data is then played back by the user to check its quality and content.

[0332] Step 7:

[0333] The user receives the audio data and plays it within the application. The user checks the sound and verifies that the audio is correctly generated in the voice selected by the voice provider.

[0334] Step 8:

[0335] The terminal calculates usage fees based on the use of voice data and sends that information back to the server. The server then collects the fees based on this information and distributes them to the voice providers as license fees.

[0336] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0337] This invention enables the generation of voices that respond to the user's emotions by integrating an emotion engine into a system that provides voices generated using voice data from voice providers to the user's terminal. The system enters into a license agreement with a voice provider and constructs a voice generation model based on that voice data.

[0338] The server receives text requests and emotion data sent from the user's terminal and first analyzes the user's emotional state using an emotion engine. The emotion engine evaluates the user's emotions in real time using user input and data obtained from the terminal's sensors. This emotion data is then input into a speech generation model, which adjusts the speech according to the user's emotions.

[0339] For example, if a user wants to generate character dialogue based on feelings of joy through an application, the text entered by the user is sent to the server. The server uses an emotion engine to evaluate the emotion data obtained from the device and incorporates the results into the voice generation process, thereby generating voice with a nuance of joy. This voice is then accompanied by official authentication information and sent to the user's device.

[0340] Users review the received audio and pay a usage fee within the application based on their usage. After the server collects the usage fee, it returns it to the audio provider as a license fee. This system allows audio providers to earn revenue based on their audio assets, and users can confidently use audio content that resonates with their emotions.

[0341] The following describes the processing flow.

[0342] Step 1:

[0343] The user launches the application and accesses the interface for speech generation. The user enters the content of the speech they want to generate as text.

[0344] Step 2:

[0345] The device acquires text data entered by the user and also collects the user's emotional data through emotion recognition sensors (camera, microphone, biometric data, etc.).

[0346] Step 3:

[0347] The device sends the collected text and emotion data to the server and requests speech generation.

[0348] Step 4:

[0349] The server uses an emotion engine to analyze the received emotion data and evaluate the user's specific emotions (joy, sadness, anger, etc.).

[0350] Step 5:

[0351] The server adjusts the parameters of the speech generation model based on the user's emotion evaluation results. This ensures that the emotion is reflected in the tone and intonation of the generated speech.

[0352] Step 6:

[0353] The server converts text data into speech data using a finely tuned speech generation model.

[0354] Step 7:

[0355] The server assigns official authentication information to the generated audio data, creating reliable and rights-protected audio data.

[0356] Step 8:

[0357] The server sends audio data containing official authentication information to the terminal and provides it to the user.

[0358] Step 9:

[0359] The user reviews the received audio data. If the user is satisfied with the content and emotional expression of the audio, they decide to pay a usage fee through the application.

[0360] Step 10:

[0361] The device sends payment information to the server and processes the usage fee.

[0362] Step 11:

[0363] The server collects usage fees from users and returns them to the voice providers as license fees in proportion to the licensing agreement.

[0364] (Example 2)

[0365] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0366] Conventional voice generation systems have difficulty generating voices that respond to user emotions, and have failed to accurately reflect the emotional nuances desired by the user. Furthermore, they lacked mechanisms for effectively utilizing voice providers' voice data and managing licenses appropriately. It is necessary to solve these problems and achieve more natural and emotionally resonant voice generation.

[0367] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0368] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data, means for analyzing the user's emotional state based on input data from a terminal, and means for generating voice data using a voice generation model in response to requests from the user terminal and the analyzed emotional state. This enables voice generation that responds to the user's emotions, and allows for the provision of voice content that meets the diverse needs of users while appropriately managing the voice provider's voice assets.

[0369] A "license agreement for providing audio data" is a legal agreement in which an audio provider grants others the right to use their audio data and receives compensation for it.

[0370] A "speech generation model" is a computational algorithm or mathematical model that takes audio data as input and generates new audio based on that data.

[0371] "Means for analyzing a user's emotional state based on input data from a device" refers to systems or methods for identifying a user's current emotions based on data (voice, video, text, etc.) obtained from the device the user is using.

[0372] A "user terminal request" refers to instructions or information sent by a user through their device to request a specific service or function.

[0373] "Means of providing official authentication information" refers to the process of adding information to generated data or content to prove its legitimacy and origin.

[0374] "Means of distributing licensing fees to audio providers" refers to methods or systems for appropriately distributing usage fees received from users to providers who hold the rights to the audio data.

[0375] This invention is a system that utilizes voice data from voice providers to generate voices that respond to the user's emotions. Specifically, a license agreement is concluded with the voice provider, and a voice generation model is built based on that voice data. The server receives text requests and emotion data from the user terminal. The emotion data is obtained from user input and terminal sensors and analyzed by the emotion engine. Based on the results of this analysis, the voice generation model generates voices that match the user's emotions.

[0376] The system implementation utilizes various hardware and software. Specifically, the server runs on a cloud platform, and speech generation uses generative AI technology and a speech synthesis API. For emotion analysis, an emotion analysis API is used to evaluate the user's emotions in real time. The user terminal inputs a voice request and runs an application to receive and play the generated voice. This allows users to easily obtain voice content that matches their desired emotion.

[0377] For example, if a user wants to generate an audio message that is "cheerful and encouraging," they would enter the following prompt.

[0378] Example of a prompt: "Let's do our best today. Please speak in a cheerful and bright voice."

[0379] Based on this input, the server performs appropriate sentiment analysis and uses a speech generation model to generate speech that reflects an "energetic tone," which is then delivered to the user's terminal. This system allows users to quickly obtain speech that matches their emotions, and enables speech providers to secure revenue by utilizing their speech assets.

[0380] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0381] Step 1:

[0382] The user uses their device to input prompt text for voice generation. For example, they might enter text such as "Please create an uplifting message" into the application. The entered text is sent from the device to the server. This data is used to verify whether the entered text aligns with the user's expected emotional response.

[0383] Step 2:

[0384] The server receives emotional data from the terminal's sensors (e.g., a facial recognition camera or voice tone microphone) along with the prompt message received from the terminal. This allows the server to understand the user's current emotional state and use it as data necessary for analysis. The server uses an emotion engine to analyze the acquired data and evaluate the user's emotional state. This evaluation result becomes the output used to adjust the speech generation model.

[0385] Step 3:

[0386] The server inputs the evaluated emotion data and prompt text into a speech generation model. This model utilizes generative AI technology to process the user's emotions and text to generate appropriate speech. The speech generation model outputs speech data that reflects the emotions, incorporating the nuances the user expects into the speech.

[0387] Step 4:

[0388] The server adds official authentication information to the generated audio data, guaranteeing that the original audio data is legitimate. This audio data is then ready to be sent to the user's terminal. The addition of official authentication information is a specific action taken to ensure the reliability of the data.

[0389] Step 5:

[0390] The user's device receives audio data sent from the server, and the user plays the audio using an application. The user can review the audio and request regeneration if necessary. This process allows the user to verify the audio quality and emotional appropriateness.

[0391] Step 6:

[0392] Users pay a fee for the use of the generated audio, and the server manages this payment information. Once the payment processing is complete, the server distributes the license fee to the audio provider. This allows the audio provider to earn revenue through their audio data. The payment and distribution stages are crucial processes for ensuring the overall profitability of the service and the convenience of the user.

[0393] (Application Example 2)

[0394] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0395] In audio content distribution, there is a growing need to provide users with a more immersive experience by offering audio that responds to their emotional state. A method for efficiently delivering services that meet this requirement is necessary. Furthermore, fair compensation for audio providers must also be considered.

[0396] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0397] In this invention, the server includes means for concluding a licensing agreement with a voice provider to provide voice data, means for creating a voice generation model based on the voice data, and means for analyzing the user's emotional state and acquiring emotional data. This makes it possible to generate and provide voice data that corresponds to the user's emotions.

[0398] A "voice provider" is a person who has the right to provide voice data and whose data is used to construct a voice generation model.

[0399] A "licensing agreement" is a contract that forms an agreement with the audio provider for the use of specific rights.

[0400] A "speech generation model" is an algorithm or program that generates speech that corresponds to the user's emotions based on the speech data of a speech provider.

[0401] An "emotion analysis device" is a device that analyzes a user's emotional state based on information provided by the user and data from terminal sensors.

[0402] "Voice data generation" is the process of generating voice data to be provided to the user using a voice generation model based on acquired emotional data.

[0403] "Official authentication information" refers to identification information assigned to prove that the generated audio data is legitimate.

[0404] "Usage fees" refer to the fees that users must pay to use the generated audio data.

[0405] "License fees" refer to the payment made to the voice provider for the right to use the voice data and build the model.

[0406] To implement this invention, a voice provider provides voice data to a server, and the server constructs a voice generation model based on that data. A licensing agreement is concluded between the voice provider and the server, securing the right to use the voice data. The server receives emotion data transmitted from the user terminal via an emotion analysis device. This emotion data is acquired through user input and sensors and is used to analyze the user's emotional state in real time.

[0407] The server analyzes the user's emotional state and uses a speech generation model to generate audio data corresponding to that emotional data. The generated audio data is accompanied by official authentication information to guarantee its authenticity. The speech generation model is developed using Python and related libraries (PyTorch, TensorFlow). The generated audio data is transmitted to the user's terminal via the internet.

[0408] The user confirms that the audio data has been provided on their device and pays the usage fee through a digital payment system. The server distributes the usage fee to the audio provider as a licensing fee. This allows users to use audio content that fits their emotions, and audio providers to receive compensation for the audio data they provide.

[0409] For example, when a user is listening to a horror novel in an audiobook application, if the emotion analyzer recognizes the user's level of tension, the voice generated by the model will be adjusted to a gentler tone. This mechanism improves the accuracy of emotion-based voice generation and enhances the user experience. The voice generation process can be controlled using a prompt such as, "Generate voice reading the next text in a manner appropriate for when the user's stress level is high."

[0410] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0411] Step 1:

[0412] The user terminal requests audio content and collects data related to the user's emotions during the process. It receives user voice input and sensor information as input, processes it with an emotion analysis device, and outputs emotion data. Specifically, it uses the user terminal's microphone and camera to capture the user's facial expressions and voice tone in real time.

[0413] Step 2:

[0414] The server receives emotion data sent from the user's terminal. Based on the input emotion data, it uses an emotion analysis engine to analyze the user's emotional state and outputs the results in a data format. Specifically, it analyzes the obtained emotion information and evaluates the user's emotional state (e.g., joy, sadness, tension).

[0415] Step 3:

[0416] The server constructs a speech generation model based on the analyzed emotion data and the user's request text. The emotion data and text are input to the speech generation model, and the adjusted speech data is output. Specifically, the generation AI model built on the server generates a speech waveform corresponding to the emotion.

[0417] Step 4:

[0418] The server assigns official authentication information to the generated audio data. It labels the audio data with authentication information in text format and outputs the result. Specifically, it adds a digital signature to the generated audio file to verify its authenticity.

[0419] Step 5:

[0420] The server sends audio data with official authentication information attached to the user's terminal. The audio data is transferred using a communication protocol, and the audio file is output to the user's terminal. Specifically, the audio data is streamed in real time over the internet.

[0421] Step 6:

[0422] The user receives and confirms the voice data on their device and pays the usage fee. They verify that the voice is as expected as input and pay the fee as output through a digital payment system. Specifically, they confirm the billing information on the application and complete the payment process.

[0423] Step 7:

[0424] The server distributes the collected royalties to the voice providers as licensing fees. Based on the royalty data received as input, it generates payment orders to the voice providers as output. Specifically, it uses an electronic payment platform to automatically process transfers to the voice providers.

[0425] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0426] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0427] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0428] [Third Embodiment]

[0429] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0430] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0431] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0432] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0433] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0435] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0436] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0437] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0438] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0439] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0440] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0441] This invention provides a system that allows voice providers to profit from the use of voice generation models based on their own voice data. This system enters into licensing agreements with voice providers, collects their voice data, and develops voice generation models. The developed models are integrated into user applications and accept voice generation requests from user terminals.

[0442] The server generates audio data using a speech generation model based on requests sent from the user's terminal. During this process, official authentication information is attached to the generated audio data, enabling the provision of reliable audio to the user. The user receives the generated audio through the application and pays a usage fee based on its use. The paid usage fees are collected by the server and distributed to the audio providers as license fees.

[0443] As a concrete example, if a user wants to generate a message in the voice of a specific character, the user launches the application and enters the desired text. The device sends this request to the server, which uses a speech generation model to create the voice. The generated voice is immediately sent to the user's device, and its quality and reliability are guaranteed by official authentication information. After the user confirms the voice, the usage fee is processed within the application, and the license fee is returned to the voice provider, who is the voice actor. This system allows voice actors to earn a sustainable income by utilizing their voices, and fans can enjoy messages using the voice actors' voices with peace of mind.

[0444] The following describes the processing flow.

[0445] Step 1:

[0446] The server will enter into a license agreement with the audio provider and define the procedure for the audio provider to provide their audio data to the server.

[0447] Step 2:

[0448] The server applies a speech processing algorithm to the received audio data in order to begin developing a speech generation model.

[0449] Step 3:

[0450] The server integrates the developed speech generation model into an application accessible to users and distributes the application.

[0451] Step 4:

[0452] Users download the application and create and configure a user account when they begin using it.

[0453] Step 5:

[0454] The user enters the text they want to have spoken in the application's interface and initiates the request.

[0455] Step 6:

[0456] The terminal sends the entered text information to the server and requests the generation of audio data.

[0457] Step 7:

[0458] The server generates speech data based on the input text using a speech generation model.

[0459] Step 8:

[0460] The server assigns official authentication information to the generated audio data, guaranteeing the reliability of the audio data.

[0461] Step 9:

[0462] The server sends voice data containing official authentication information to the terminal, making it available to the user.

[0463] Step 10:

[0464] The user reviews the received audio data and then proceeds to pay the usage fee for its use.

[0465] Step 11:

[0466] The device sends payment information for the usage fee to the server, and the payment is confirmed.

[0467] Step 12:

[0468] The server collects usage fees and distributes a portion of them to the voice providers as licensing fees, thereby returning profits to them.

[0469] (Example 1)

[0470] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0471] Conventional voice generation systems have problems with adequately protecting the rights of voice providers and with users being unable to obtain reliable audio information. Furthermore, there is a lack of mechanisms to guarantee the reliability and quality of the voice generation models used, making it difficult to provide reliable value to users. In addition, mechanisms for efficiently managing and fairly distributing royalties to rights holders when users utilize specific voices are immature.

[0472] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0473] In this invention, the server includes means for concluding a rights agreement with the speaker to permit the provision of acoustic information, means for creating a speech generation algorithm based on the speaker's acoustic information, and means for generating acoustic information using the speech generation algorithm in response to a request from a user device. This ensures the reliability and quality of the acoustic information through official authentication information, enabling reliable speech generation for users and protecting the rights of the speaker, as well as effective management and distribution of royalties.

[0474] "Acoustic information" refers to data that digitally represents the characteristics of sound propagating through space.

[0475] A "rights agreement" refers to a legal agreement in which an audio provider grants permission for a third party to use their audio information.

[0476] A "speaker" refers to someone who provides their own sonic information.

[0477] A "speech generation algorithm" refers to a computational method or process for generating new acoustic information based on existing acoustic information.

[0478] A "user device" refers to an electronic device that a user can directly operate and that has the function of generating, receiving, and playing back acoustic information.

[0479] A "request" refers to a command sent from a user device to a server requesting specific processing or information provision.

[0480] "Official authentication information" refers to the digital signatures and metadata attached to prove that the generated audio information is trustworthy.

[0481] "Presentation text" refers to the text data input into a speech generation algorithm, and is a sentence that indicates the content that forms the basis of the generated acoustic information.

[0482] An "input / output mechanism" refers to an interface that allows a user to input or output data through their device.

[0483] "Transmission" refers to the process of sending generated acoustic information to the user's device.

[0484] This invention is realized by a speech generation system based on the acoustic information of voice providers. The server enters into a rights agreement with the voice providers and collects their acoustic information. The collected acoustic information is used to develop a speech generation algorithm. This algorithm is trained using machine learning frameworks such as TensorFlow and PyTorch and built as a model capable of generating new acoustic information.

[0485] If a user wants to generate audio information, they use an application installed on their device. The user enters the desired text into this application. Specifically, they might enter a prompt such as, "Convert the following text into speech in the voice of the specified character: 'Good morning! What fun things await me today?'" This entered data is then sent from the device to the server.

[0486] The server inputs the received text data into a configured speech generation algorithm and uses a generation AI model to generate audio information. During this process, the server guarantees the reliability and quality of the information by attaching official authentication information to the generated audio. Finally, the generated audio information is sent from the server to the user's terminal, where the user can review and use the audio.

[0487] Through this invention, audio providers can earn revenue while protecting their rights, and users can easily access high-quality, reliable audio information.

[0488] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0489] Step 1:

[0490] The server enters into a rights agreement with the voice provider and collects acoustic information. The input for this step is the acoustic information provided by the voice provider. The server stores the acoustic information as digital data in preparation for later model development.

[0491] Step 2:

[0492] The server uses the collected acoustic information to train a speech generation algorithm. The input for this step is the stored acoustic information as digital data. The server leverages a machine learning platform (e.g., TensorFlow) to train a generative AI model. The output is a generative AI model equipped with the speech generation algorithm.

[0493] Step 3:

[0494] The user enters the desired text into the application on their device. This input is text data based on the user's preferences. The device uses this data as a prompt and sends it to the server as a speech generation request.

[0495] Step 4:

[0496] The server provides the received text data to the AI ​​model as a prompt. The input for this step is the text data sent by the user. The server uses a speech generation algorithm to generate new acoustic information based on this text. The output is the generated acoustic information.

[0497] Step 5:

[0498] The server adds official authentication information to the generated acoustic information. This step adds metadata indicating the reliability of the acoustic information. The input to this process is the acoustic information resulting from speech generation, and the output is the acoustic information with official authentication information added.

[0499] Step 6:

[0500] The server sends audio information with official authentication credentials to the user's device. The input in this step is authenticated audio information. The server uses real-time communication technology to stream the data, enabling audio playback on the device.

[0501] Step 7:

[0502] The user confirms the audio information received on the device. After confirmation, the user pays the usage fee. The input is the played audio information and the user's decision to use it. The device sends the payment information to the server, and the output is the successfully processed payment data.

[0503] Step 8:

[0504] The server distributes the collected royalties to the voice providers. The input for this step is payment data from the user. Based on the contract with the voice provider, the server calculates the royalties, and the output is the distribution of appropriate license fees to the voice provider.

[0505] (Application Example 1)

[0506] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0507] The challenge is to provide a system that delivers reliable audio content to audio content enthusiasts and enables audio providers to continuously generate revenue. Another challenge is to enable users to instantly generate and obtain audio information tailored to their preferences.

[0508] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0509] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data; means for creating a voice generation model based on the voice provider's voice data; means for generating acoustic information using the voice generation model in response to a request from a user terminal; means for attaching official authentication information to the generated acoustic information; means for collecting usage fees from users and distributing them to voice providers as license fees; means for providing an information input area for generating acoustic information; and means for transmitting the generated acoustic information to an information terminal. As a result, users can obtain reliable acoustic content with peace of mind, and voice providers can achieve stable monetization.

[0510] A "voice provider" is an individual or organization that provides their own voice data and cooperates in the development of a voice generation model.

[0511] A "license agreement" is a legal agreement entered into by an audio provider to permit a third party to use the audio data.

[0512] A "speech generation model" is an algorithm or program built on speech data from a speech provider to generate speech corresponding to specified text.

[0513] A "user terminal" is an electronic device owned by a user and used to generate and receive acoustic information.

[0514] "Official authentication information" refers to digital information indicating reliability, which is assigned to guarantee that the generated audio information is genuine.

[0515] A "usage fee" is money paid by a user as compensation for acquiring and using audio information.

[0516] "Acoustic information" refers to audio data created using a speech generation model.

[0517] The "information input area" is the interface section provided for users to input the text and parameters necessary for generating acoustic information.

[0518] An "information terminal" is a digital device used to receive and play back generated audio information.

[0519] This invention is a system that develops a speech generation model using speech data from a speech provider and provides reliable acoustic information to the user. The server comprises several key components for this purpose. It creates a speech generation model based on speech data obtained from the speech provider and operates efficiently using cloud-based computing resources. This model is built using a deep learning framework such as TensorFlow. The server also leverages Google Cloud's Text to Speech API to attach official authentication information to the generated acoustic information.

[0520] The user terminal refers to an information device such as a smartphone or tablet, which sends a voice generation request through an application. The user can enter text into an information input area within the application and select their preferred voice provider. This request is sent to a server via the internet, where the necessary audio information is generated by a designated generation AI model. The generated audio information is returned to the user terminal, where its quality is verified by official authentication information. Users can then enjoy the audio content with peace of mind.

[0521] A concrete example of this use case is when a user creates a message to a friend using the voice of their favorite voice actor. The user enters a text message such as "Hello, I'm the AI ​​assistant. Have a great day!" into the application's text input area, and the application generates a message in the voice of that voice actor. Based on this prompt, the server responds immediately and provides reliable audio information to the user's device.

[0522] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0523] Step 1:

[0524] The user launches the application on their smartphone or tablet and enters the text of the voice message they want to generate into the information input area. This entered text is the input data used in the next step of the process.

[0525] Step 2:

[0526] The terminal sends the voice provider and text information to the server according to the user's selection. This information is processed on the server side as basic data for selecting a speech generation model and for text-to-speech conversion.

[0527] Step 3:

[0528] The server selects the optimal speech generation model based on the received text and information from the voice provider. This model operates using the TensorFlow framework and generates acoustic information in the specified voice using a deep learning algorithm.

[0529] Step 4:

[0530] The selected speech generation model uses Google Cloud's Text to Speech API to convert text into speech data. This process involves parsing the text data and generating the corresponding speech waveform. The converted speech data is the output.

[0531] Step 5:

[0532] The server assigns official authentication information to the generated audio data. This information is generated using digital signatures and other methods to guarantee the authenticity of the audio.

[0533] Step 6:

[0534] The server sends authenticated audio data to the user's device. This data is then played back by the user to check its quality and content.

[0535] Step 7:

[0536] The user receives the audio data and plays it within the application. The user checks the sound and verifies that the audio is correctly generated in the voice selected by the voice provider.

[0537] Step 8:

[0538] The terminal calculates usage fees based on the use of voice data and sends that information back to the server. The server then collects the fees based on this information and distributes them to the voice providers as license fees.

[0539] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0540] This invention enables the generation of voices that respond to the user's emotions by integrating an emotion engine into a system that provides voices generated using voice data from voice providers to the user's terminal. The system enters into a license agreement with a voice provider and constructs a voice generation model based on that voice data.

[0541] The server receives text requests and emotion data sent from the user's terminal and first analyzes the user's emotional state using an emotion engine. The emotion engine evaluates the user's emotions in real time using user input and data obtained from the terminal's sensors. This emotion data is then input into a speech generation model, which adjusts the speech according to the user's emotions.

[0542] For example, if a user wants to generate character dialogue based on feelings of joy through an application, the text entered by the user is sent to the server. The server uses an emotion engine to evaluate the emotion data obtained from the device and incorporates the results into the voice generation process, thereby generating voice with a nuance of joy. This voice is then accompanied by official authentication information and sent to the user's device.

[0543] Users review the received audio and pay a usage fee within the application based on their usage. After the server collects the usage fee, it returns it to the audio provider as a license fee. This system allows audio providers to earn revenue based on their audio assets, and users can confidently use audio content that resonates with their emotions.

[0544] The following describes the processing flow.

[0545] Step 1:

[0546] The user launches the application and accesses the interface for speech generation. The user enters the content of the speech they want to generate as text.

[0547] Step 2:

[0548] The device acquires text data entered by the user and also collects the user's emotional data through emotion recognition sensors (camera, microphone, biometric data, etc.).

[0549] Step 3:

[0550] The device sends the collected text and emotion data to the server and requests speech generation.

[0551] Step 4:

[0552] The server uses an emotion engine to analyze the received emotion data and evaluate the user's specific emotions (joy, sadness, anger, etc.).

[0553] Step 5:

[0554] The server adjusts the parameters of the speech generation model based on the user's emotion evaluation results. This ensures that the emotion is reflected in the tone and intonation of the generated speech.

[0555] Step 6:

[0556] The server converts text data into speech data using a finely tuned speech generation model.

[0557] Step 7:

[0558] The server assigns official authentication information to the generated audio data, creating reliable and rights-protected audio data.

[0559] Step 8:

[0560] The server sends audio data containing official authentication information to the terminal and provides it to the user.

[0561] Step 9:

[0562] The user reviews the received audio data. If the user is satisfied with the content and emotional expression of the audio, they decide to pay a usage fee through the application.

[0563] Step 10:

[0564] The device sends payment information to the server and processes the usage fee.

[0565] Step 11:

[0566] The server collects usage fees from users and returns them to the voice providers as license fees in proportion to the licensing agreement.

[0567] (Example 2)

[0568] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0569] Conventional voice generation systems have difficulty generating voices that respond to user emotions, and have failed to accurately reflect the emotional nuances desired by the user. Furthermore, they lacked mechanisms for effectively utilizing voice providers' voice data and managing licenses appropriately. It is necessary to solve these problems and achieve more natural and emotionally resonant voice generation.

[0570] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0571] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data, means for analyzing the user's emotional state based on input data from a terminal, and means for generating voice data using a voice generation model in response to requests from the user terminal and the analyzed emotional state. This enables voice generation that responds to the user's emotions, and allows for the provision of voice content that meets the diverse needs of users while appropriately managing the voice provider's voice assets.

[0572] A "license agreement for providing audio data" is a legal agreement in which an audio provider grants others the right to use their audio data and receives compensation for it.

[0573] A "speech generation model" is a computational algorithm or mathematical model that takes audio data as input and generates new audio based on that data.

[0574] "Means for analyzing a user's emotional state based on input data from a device" refers to systems or methods for identifying a user's current emotions based on data (voice, video, text, etc.) obtained from the device the user is using.

[0575] A "user terminal request" refers to instructions or information sent by a user through their device to request a specific service or function.

[0576] "Means of providing official authentication information" refers to the process of adding information to generated data or content to prove its legitimacy and origin.

[0577] "Means of distributing licensing fees to audio providers" refers to methods or systems for appropriately distributing usage fees received from users to providers who hold the rights to the audio data.

[0578] This invention is a system that utilizes voice data from voice providers to generate voices that respond to the user's emotions. Specifically, a license agreement is concluded with the voice provider, and a voice generation model is built based on that voice data. The server receives text requests and emotion data from the user terminal. The emotion data is obtained from user input and terminal sensors and analyzed by the emotion engine. Based on the results of this analysis, the voice generation model generates voices that match the user's emotions.

[0579] The system implementation utilizes various hardware and software. Specifically, the server runs on a cloud platform, and speech generation uses generative AI technology and a speech synthesis API. For emotion analysis, an emotion analysis API is used to evaluate the user's emotions in real time. The user terminal inputs a voice request and runs an application to receive and play the generated voice. This allows users to easily obtain voice content that matches their desired emotion.

[0580] For example, if a user wants to generate an audio message that is "cheerful and encouraging," they would enter the following prompt.

[0581] Example of a prompt: "Let's do our best today. Please speak in a cheerful and bright voice."

[0582] Based on this input, the server performs appropriate sentiment analysis and uses a speech generation model to generate speech that reflects an "energetic tone," which is then delivered to the user's terminal. This system allows users to quickly obtain speech that matches their emotions, and enables speech providers to secure revenue by utilizing their speech assets.

[0583] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0584] Step 1:

[0585] The user uses their device to input prompt text for voice generation. For example, they might enter text such as "Please create an uplifting message" into the application. The entered text is sent from the device to the server. This data is used to verify whether the entered text aligns with the user's expected emotional response.

[0586] Step 2:

[0587] The server receives emotional data from the terminal's sensors (e.g., a facial recognition camera or voice tone microphone) along with the prompt message received from the terminal. This allows the server to understand the user's current emotional state and use it as data necessary for analysis. The server uses an emotion engine to analyze the acquired data and evaluate the user's emotional state. This evaluation result becomes the output used to adjust the speech generation model.

[0588] Step 3:

[0589] The server inputs the evaluated emotion data and prompt text into a speech generation model. This model utilizes generative AI technology to process the user's emotions and text to generate appropriate speech. The speech generation model outputs speech data that reflects the emotions, incorporating the nuances the user expects into the speech.

[0590] Step 4:

[0591] The server adds official authentication information to the generated audio data, guaranteeing that the original audio data is legitimate. This audio data is then ready to be sent to the user's terminal. The addition of official authentication information is a specific action taken to ensure the reliability of the data.

[0592] Step 5:

[0593] The user's device receives audio data sent from the server, and the user plays the audio using an application. The user can review the audio and request regeneration if necessary. This process allows the user to verify the audio quality and emotional appropriateness.

[0594] Step 6:

[0595] Users pay a fee for the use of the generated audio, and the server manages this payment information. Once the payment processing is complete, the server distributes the license fee to the audio provider. This allows the audio provider to earn revenue through their audio data. The payment and distribution stages are crucial processes for ensuring the overall profitability of the service and the convenience of the user.

[0596] (Application Example 2)

[0597] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0598] In audio content distribution, there is a growing need to provide users with a more immersive experience by offering audio that responds to their emotional state. A method for efficiently delivering services that meet this requirement is necessary. Furthermore, fair compensation for audio providers must also be considered.

[0599] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0600] In this invention, the server includes means for concluding a licensing agreement with a voice provider to provide voice data, means for creating a voice generation model based on the voice data, and means for analyzing the user's emotional state and acquiring emotional data. This makes it possible to generate and provide voice data that corresponds to the user's emotions.

[0601] A "voice provider" is a person who has the right to provide voice data and whose data is used to construct a voice generation model.

[0602] A "licensing agreement" is a contract that forms an agreement with the audio provider for the use of specific rights.

[0603] A "speech generation model" is an algorithm or program that generates speech that corresponds to the user's emotions based on the speech data of a speech provider.

[0604] An "emotion analysis device" is a device that analyzes a user's emotional state based on information provided by the user and data from terminal sensors.

[0605] "Voice data generation" is the process of generating voice data to be provided to the user using a voice generation model based on acquired emotional data.

[0606] "Official authentication information" refers to identification information assigned to prove that the generated audio data is legitimate.

[0607] "Usage fees" refer to the fees that users must pay to use the generated audio data.

[0608] "License fees" refer to the payment made to the voice provider for the right to use the voice data and build the model.

[0609] To implement this invention, a voice provider provides voice data to a server, and the server constructs a voice generation model based on that data. A licensing agreement is concluded between the voice provider and the server, securing the right to use the voice data. The server receives emotion data transmitted from the user terminal via an emotion analysis device. This emotion data is acquired through user input and sensors and is used to analyze the user's emotional state in real time.

[0610] The server analyzes the user's emotional state and uses a speech generation model to generate audio data corresponding to that emotional data. The generated audio data is accompanied by official authentication information to guarantee its authenticity. The speech generation model is developed using Python and related libraries (PyTorch, TensorFlow). The generated audio data is transmitted to the user's terminal via the internet.

[0611] The user confirms that the audio data has been provided on their device and pays the usage fee through a digital payment system. The server distributes the usage fee to the audio provider as a licensing fee. This allows users to use audio content that fits their emotions, and audio providers to receive compensation for the audio data they provide.

[0612] For example, when a user is listening to a horror novel in an audiobook application, if the emotion analyzer recognizes the user's level of tension, the voice generated by the model will be adjusted to a gentler tone. This mechanism improves the accuracy of emotion-based voice generation and enhances the user experience. The voice generation process can be controlled using a prompt such as, "Generate voice reading the next text in a manner appropriate for when the user's stress level is high."

[0613] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0614] Step 1:

[0615] The user terminal requests audio content and collects data related to the user's emotions during the process. It receives user voice input and sensor information as input, processes it with an emotion analysis device, and outputs emotion data. Specifically, it uses the user terminal's microphone and camera to capture the user's facial expressions and voice tone in real time.

[0616] Step 2:

[0617] The server receives emotion data sent from the user's terminal. Based on the input emotion data, it uses an emotion analysis engine to analyze the user's emotional state and outputs the results in a data format. Specifically, it analyzes the obtained emotion information and evaluates the user's emotional state (e.g., joy, sadness, tension).

[0618] Step 3:

[0619] The server constructs a speech generation model based on the analyzed emotion data and the user's request text. The emotion data and text are input to the speech generation model, and the adjusted speech data is output. Specifically, the generation AI model built on the server generates a speech waveform corresponding to the emotion.

[0620] Step 4:

[0621] The server assigns official authentication information to the generated audio data. It labels the audio data with authentication information in text format and outputs the result. Specifically, it adds a digital signature to the generated audio file to verify its authenticity.

[0622] Step 5:

[0623] The server sends audio data with official authentication information attached to the user's terminal. The audio data is transferred using a communication protocol, and the audio file is output to the user's terminal. Specifically, the audio data is streamed in real time over the internet.

[0624] Step 6:

[0625] The user receives and confirms the voice data on their device and pays the usage fee. They verify that the voice is as expected as input and pay the fee as output through a digital payment system. Specifically, they confirm the billing information on the application and complete the payment process.

[0626] Step 7:

[0627] The server distributes the collected royalties to the voice providers as licensing fees. Based on the royalty data received as input, it generates payment orders to the voice providers as output. Specifically, it uses an electronic payment platform to automatically process transfers to the voice providers.

[0628] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0629] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0630] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0631] [Fourth Embodiment]

[0632] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0633] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0634] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0635] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0636] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0637] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0638] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0639] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0640] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0641] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0642] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0643] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0644] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] This invention provides a system that allows voice providers to profit from the use of voice generation models based on their own voice data. This system enters into licensing agreements with voice providers, collects their voice data, and develops voice generation models. The developed models are integrated into user applications and accept voice generation requests from user terminals.

[0646] The server generates audio data using a speech generation model based on requests sent from the user's terminal. During this process, official authentication information is attached to the generated audio data, enabling the provision of reliable audio to the user. The user receives the generated audio through the application and pays a usage fee based on its use. The paid usage fees are collected by the server and distributed to the audio providers as license fees.

[0647] As a concrete example, if a user wants to generate a message in the voice of a specific character, the user launches the application and enters the desired text. The device sends this request to the server, which uses a speech generation model to create the voice. The generated voice is immediately sent to the user's device, and its quality and reliability are guaranteed by official authentication information. After the user confirms the voice, the usage fee is processed within the application, and the license fee is returned to the voice provider, who is the voice actor. This system allows voice actors to earn a sustainable income by utilizing their voices, and fans can enjoy messages using the voice actors' voices with peace of mind.

[0648] The following describes the processing flow.

[0649] Step 1:

[0650] The server will enter into a license agreement with the audio provider and define the procedure for the audio provider to provide their audio data to the server.

[0651] Step 2:

[0652] The server applies a speech processing algorithm to the received audio data in order to begin developing a speech generation model.

[0653] Step 3:

[0654] The server integrates the developed speech generation model into an application accessible to users and distributes the application.

[0655] Step 4:

[0656] Users download the application and create and configure a user account when they begin using it.

[0657] Step 5:

[0658] The user enters the text they want to have spoken in the application's interface and initiates the request.

[0659] Step 6:

[0660] The terminal sends the entered text information to the server and requests the generation of audio data.

[0661] Step 7:

[0662] The server generates speech data based on the input text using a speech generation model.

[0663] Step 8:

[0664] The server assigns official authentication information to the generated audio data, guaranteeing the reliability of the audio data.

[0665] Step 9:

[0666] The server sends voice data containing official authentication information to the terminal, making it available to the user.

[0667] Step 10:

[0668] The user reviews the received audio data and then proceeds to pay the usage fee for its use.

[0669] Step 11:

[0670] The device sends payment information for the usage fee to the server, and the payment is confirmed.

[0671] Step 12:

[0672] The server collects usage fees and distributes a portion of them to the voice providers as licensing fees, thereby returning profits to them.

[0673] (Example 1)

[0674] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0675] Conventional voice generation systems have problems with adequately protecting the rights of voice providers and with users being unable to obtain reliable audio information. Furthermore, there is a lack of mechanisms to guarantee the reliability and quality of the voice generation models used, making it difficult to provide reliable value to users. In addition, mechanisms for efficiently managing and fairly distributing royalties to rights holders when users utilize specific voices are immature.

[0676] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0677] In this invention, the server includes means for concluding a rights agreement with the speaker to permit the provision of acoustic information, means for creating a speech generation algorithm based on the speaker's acoustic information, and means for generating acoustic information using the speech generation algorithm in response to a request from a user device. This ensures the reliability and quality of the acoustic information through official authentication information, enabling reliable speech generation for users and protecting the rights of the speaker, as well as effective management and distribution of royalties.

[0678] "Acoustic information" refers to data that digitally represents the characteristics of sound propagating through space.

[0679] A "rights agreement" refers to a legal agreement in which an audio provider grants permission for a third party to use their audio information.

[0680] A "speaker" refers to someone who provides their own sonic information.

[0681] A "speech generation algorithm" refers to a computational method or process for generating new acoustic information based on existing acoustic information.

[0682] A "user device" refers to an electronic device that a user can directly operate and that has the function of generating, receiving, and playing back acoustic information.

[0683] A "request" refers to a command sent from a user device to a server requesting specific processing or information provision.

[0684] "Official authentication information" refers to the digital signatures and metadata attached to prove that the generated audio information is trustworthy.

[0685] "Presentation text" refers to the text data input into a speech generation algorithm, and is a sentence that indicates the content that forms the basis of the generated acoustic information.

[0686] An "input / output mechanism" refers to an interface that allows a user to input or output data through their device.

[0687] "Transmission" refers to the process of sending generated acoustic information to the user's device.

[0688] This invention is realized by a speech generation system based on the acoustic information of voice providers. The server enters into a rights agreement with the voice providers and collects their acoustic information. The collected acoustic information is used to develop a speech generation algorithm. This algorithm is trained using machine learning frameworks such as TensorFlow and PyTorch and built as a model capable of generating new acoustic information.

[0689] If a user wants to generate audio information, they use an application installed on their device. The user enters the desired text into this application. Specifically, they might enter a prompt such as, "Convert the following text into speech in the voice of the specified character: 'Good morning! What fun things await me today?'" This entered data is then sent from the device to the server.

[0690] The server inputs the received text data into a configured speech generation algorithm and uses a generation AI model to generate audio information. During this process, the server guarantees the reliability and quality of the information by attaching official authentication information to the generated audio. Finally, the generated audio information is sent from the server to the user's terminal, where the user can review and use the audio.

[0691] Through this invention, audio providers can earn revenue while protecting their rights, and users can easily access high-quality, reliable audio information.

[0692] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0693] Step 1:

[0694] The server enters into a rights agreement with the voice provider and collects acoustic information. The input for this step is the acoustic information provided by the voice provider. The server stores the acoustic information as digital data in preparation for later model development.

[0695] Step 2:

[0696] The server uses the collected acoustic information to train a speech generation algorithm. The input for this step is the stored acoustic information as digital data. The server leverages a machine learning platform (e.g., TensorFlow) to train a generative AI model. The output is a generative AI model equipped with the speech generation algorithm.

[0697] Step 3:

[0698] The user enters the desired text into the application on their device. This input is text data based on the user's preferences. The device uses this data as a prompt and sends it to the server as a speech generation request.

[0699] Step 4:

[0700] The server provides the received text data to the AI ​​model as a prompt. The input for this step is the text data sent by the user. The server uses a speech generation algorithm to generate new acoustic information based on this text. The output is the generated acoustic information.

[0701] Step 5:

[0702] The server adds official authentication information to the generated acoustic information. This step adds metadata indicating the reliability of the acoustic information. The input to this process is the acoustic information resulting from speech generation, and the output is the acoustic information with official authentication information added.

[0703] Step 6:

[0704] The server sends audio information with official authentication credentials to the user's device. The input in this step is authenticated audio information. The server uses real-time communication technology to stream the data, enabling audio playback on the device.

[0705] Step 7:

[0706] The user confirms the audio information received on the device. After confirmation, the user pays the usage fee. The input is the played audio information and the user's decision to use it. The device sends the payment information to the server, and the output is the successfully processed payment data.

[0707] Step 8:

[0708] The server distributes the collected royalties to the voice providers. The input for this step is payment data from the user. Based on the contract with the voice provider, the server calculates the royalties, and the output is the distribution of appropriate license fees to the voice provider.

[0709] (Application Example 1)

[0710] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0711] The challenge is to provide a system that delivers reliable audio content to audio content enthusiasts and enables audio providers to continuously generate revenue. Another challenge is to enable users to instantly generate and obtain audio information tailored to their preferences.

[0712] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0713] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data; means for creating a voice generation model based on the voice provider's voice data; means for generating acoustic information using the voice generation model in response to a request from a user terminal; means for attaching official authentication information to the generated acoustic information; means for collecting usage fees from users and distributing them to voice providers as license fees; means for providing an information input area for generating acoustic information; and means for transmitting the generated acoustic information to an information terminal. As a result, users can obtain reliable acoustic content with peace of mind, and voice providers can achieve stable monetization.

[0714] A "voice provider" is an individual or organization that provides their own voice data and cooperates in the development of a voice generation model.

[0715] A "license agreement" is a legal agreement entered into by an audio provider to permit a third party to use the audio data.

[0716] A "speech generation model" is an algorithm or program built on speech data from a speech provider to generate speech corresponding to specified text.

[0717] A "user terminal" is an electronic device owned by a user and used to generate and receive acoustic information.

[0718] "Official authentication information" refers to digital information indicating reliability, which is assigned to guarantee that the generated audio information is genuine.

[0719] A "usage fee" is money paid by a user as compensation for acquiring and using audio information.

[0720] "Acoustic information" refers to audio data created using a speech generation model.

[0721] The "information input area" is the interface section provided for users to input the text and parameters necessary for generating acoustic information.

[0722] An "information terminal" is a digital device used to receive and play back generated audio information.

[0723] This invention is a system that develops a speech generation model using speech data from a speech provider and provides reliable acoustic information to the user. The server comprises several key components for this purpose. It creates a speech generation model based on speech data obtained from the speech provider and operates efficiently using cloud-based computing resources. This model is built using a deep learning framework such as TensorFlow. The server also leverages Google Cloud's Text to Speech API to attach official authentication information to the generated acoustic information.

[0724] The user terminal refers to an information device such as a smartphone or tablet, which sends a voice generation request through an application. The user can enter text into an information input area within the application and select their preferred voice provider. This request is sent to a server via the internet, where the necessary audio information is generated by a designated generation AI model. The generated audio information is returned to the user terminal, where its quality is verified by official authentication information. Users can then enjoy the audio content with peace of mind.

[0725] A concrete example of this use case is when a user creates a message to a friend using the voice of their favorite voice actor. The user enters a text message such as "Hello, I'm the AI ​​assistant. Have a great day!" into the application's text input area, and the application generates a message in the voice of that voice actor. Based on this prompt, the server responds immediately and provides reliable audio information to the user's device.

[0726] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0727] Step 1:

[0728] The user launches the application on their smartphone or tablet and enters the text of the voice message they want to generate into the information input area. This entered text is the input data used in the next step of the process.

[0729] Step 2:

[0730] The terminal sends the voice provider and text information to the server according to the user's selection. This information is processed on the server side as basic data for selecting a speech generation model and for text-to-speech conversion.

[0731] Step 3:

[0732] The server selects the optimal speech generation model based on the received text and information from the voice provider. This model operates using the TensorFlow framework and generates acoustic information in the specified voice using a deep learning algorithm.

[0733] Step 4:

[0734] The selected speech generation model uses Google Cloud's Text to Speech API to convert text into speech data. This process involves parsing the text data and generating the corresponding speech waveform. The converted speech data is the output.

[0735] Step 5:

[0736] The server assigns official authentication information to the generated audio data. This information is generated using digital signatures and other methods to guarantee the authenticity of the audio.

[0737] Step 6:

[0738] The server sends authenticated audio data to the user's device. This data is then played back by the user to check its quality and content.

[0739] Step 7:

[0740] The user receives the audio data and plays it within the application. The user checks the sound and verifies that the audio is correctly generated in the voice selected by the voice provider.

[0741] Step 8:

[0742] The terminal calculates usage fees based on the use of voice data and sends that information back to the server. The server then collects the fees based on this information and distributes them to the voice providers as license fees.

[0743] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0744] This invention enables the generation of voices that respond to the user's emotions by integrating an emotion engine into a system that provides voices generated using voice data from voice providers to the user's terminal. The system enters into a license agreement with a voice provider and constructs a voice generation model based on that voice data.

[0745] The server receives text requests and emotion data sent from the user's terminal and first analyzes the user's emotional state using an emotion engine. The emotion engine evaluates the user's emotions in real time using user input and data obtained from the terminal's sensors. This emotion data is then input into a speech generation model, which adjusts the speech according to the user's emotions.

[0746] For example, if a user wants to generate character dialogue based on feelings of joy through an application, the text entered by the user is sent to the server. The server uses an emotion engine to evaluate the emotion data obtained from the device and incorporates the results into the voice generation process, thereby generating voice with a nuance of joy. This voice is then accompanied by official authentication information and sent to the user's device.

[0747] Users review the received audio and pay a usage fee within the application based on their usage. After the server collects the usage fee, it returns it to the audio provider as a license fee. This system allows audio providers to earn revenue based on their audio assets, and users can confidently use audio content that resonates with their emotions.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] The user launches the application and accesses the interface for speech generation. The user enters the content of the speech they want to generate as text.

[0751] Step 2:

[0752] The device acquires text data entered by the user and also collects the user's emotional data through emotion recognition sensors (camera, microphone, biometric data, etc.).

[0753] Step 3:

[0754] The device sends the collected text and emotion data to the server and requests speech generation.

[0755] Step 4:

[0756] The server uses an emotion engine to analyze the received emotion data and evaluate the user's specific emotions (joy, sadness, anger, etc.).

[0757] Step 5:

[0758] The server adjusts the parameters of the speech generation model based on the user's emotion evaluation results. This ensures that the emotion is reflected in the tone and intonation of the generated speech.

[0759] Step 6:

[0760] The server converts text data into speech data using a finely tuned speech generation model.

[0761] Step 7:

[0762] The server assigns official authentication information to the generated audio data, creating reliable and rights-protected audio data.

[0763] Step 8:

[0764] The server sends audio data containing official authentication information to the terminal and provides it to the user.

[0765] Step 9:

[0766] The user reviews the received audio data. If the user is satisfied with the content and emotional expression of the audio, they decide to pay a usage fee through the application.

[0767] Step 10:

[0768] The device sends payment information to the server and processes the usage fee.

[0769] Step 11:

[0770] The server collects usage fees from users and returns them to the voice providers as license fees in proportion to the licensing agreement.

[0771] (Example 2)

[0772] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0773] Conventional voice generation systems have difficulty generating voices that respond to user emotions, and have failed to accurately reflect the emotional nuances desired by the user. Furthermore, they lacked mechanisms for effectively utilizing voice providers' voice data and managing licenses appropriately. It is necessary to solve these problems and achieve more natural and emotionally resonant voice generation.

[0774] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0775] In this invention, the server includes means for concluding a license agreement with a voice provider to permit the provision of voice data, means for analyzing the user's emotional state based on input data from a terminal, and means for generating voice data using a voice generation model in response to requests from the user terminal and the analyzed emotional state. This enables voice generation that responds to the user's emotions, and allows for the provision of voice content that meets the diverse needs of users while appropriately managing the voice provider's voice assets.

[0776] A "license agreement for providing audio data" is a legal agreement in which an audio provider grants others the right to use their audio data and receives compensation for it.

[0777] A "speech generation model" is a computational algorithm or mathematical model that takes audio data as input and generates new audio based on that data.

[0778] "Means for analyzing a user's emotional state based on input data from a device" refers to systems or methods for identifying a user's current emotions based on data (voice, video, text, etc.) obtained from the device the user is using.

[0779] A "user terminal request" refers to instructions or information sent by a user through their device to request a specific service or function.

[0780] "Means of providing official authentication information" refers to the process of adding information to generated data or content to prove its legitimacy and origin.

[0781] "Means of distributing licensing fees to audio providers" refers to methods or systems for appropriately distributing usage fees received from users to providers who hold the rights to the audio data.

[0782] This invention is a system that utilizes voice data from voice providers to generate voices that respond to the user's emotions. Specifically, a license agreement is concluded with the voice provider, and a voice generation model is built based on that voice data. The server receives text requests and emotion data from the user terminal. The emotion data is obtained from user input and terminal sensors and analyzed by the emotion engine. Based on the results of this analysis, the voice generation model generates voices that match the user's emotions.

[0783] The system implementation utilizes various hardware and software. Specifically, the server runs on a cloud platform, and speech generation uses generative AI technology and a speech synthesis API. For emotion analysis, an emotion analysis API is used to evaluate the user's emotions in real time. The user terminal inputs a voice request and runs an application to receive and play the generated voice. This allows users to easily obtain voice content that matches their desired emotion.

[0784] For example, if a user wants to generate an audio message that is "cheerful and encouraging," they would enter the following prompt.

[0785] Example of a prompt: "Let's do our best today. Please speak in a cheerful and bright voice."

[0786] Based on this input, the server performs appropriate sentiment analysis and uses a speech generation model to generate speech that reflects an "energetic tone," which is then delivered to the user's terminal. This system allows users to quickly obtain speech that matches their emotions, and enables speech providers to secure revenue by utilizing their speech assets.

[0787] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0788] Step 1:

[0789] The user uses their device to input prompt text for voice generation. For example, they might enter text such as "Please create an uplifting message" into the application. The entered text is sent from the device to the server. This data is used to verify whether the entered text aligns with the user's expected emotional response.

[0790] Step 2:

[0791] The server receives emotional data from the terminal's sensors (e.g., a facial recognition camera or voice tone microphone) along with the prompt message received from the terminal. This allows the server to understand the user's current emotional state and use it as data necessary for analysis. The server uses an emotion engine to analyze the acquired data and evaluate the user's emotional state. This evaluation result becomes the output used to adjust the speech generation model.

[0792] Step 3:

[0793] The server inputs the evaluated emotion data and prompt text into a speech generation model. This model utilizes generative AI technology to process the user's emotions and text to generate appropriate speech. The speech generation model outputs speech data that reflects the emotions, incorporating the nuances the user expects into the speech.

[0794] Step 4:

[0795] The server adds official authentication information to the generated audio data, guaranteeing that the original audio data is legitimate. This audio data is then ready to be sent to the user's terminal. The addition of official authentication information is a specific action taken to ensure the reliability of the data.

[0796] Step 5:

[0797] The user's device receives audio data sent from the server, and the user plays the audio using an application. The user can review the audio and request regeneration if necessary. This process allows the user to verify the audio quality and emotional appropriateness.

[0798] Step 6:

[0799] Users pay a fee for the use of the generated audio, and the server manages this payment information. Once the payment processing is complete, the server distributes the license fee to the audio provider. This allows the audio provider to earn revenue through their audio data. The payment and distribution stages are crucial processes for ensuring the overall profitability of the service and the convenience of the user.

[0800] (Application Example 2)

[0801] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0802] In audio content distribution, there is a growing need to provide users with a more immersive experience by offering audio that responds to their emotional state. A method for efficiently delivering services that meet this requirement is necessary. Furthermore, fair compensation for audio providers must also be considered.

[0803] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0804] In this invention, the server includes means for concluding a licensing agreement with a voice provider to provide voice data, means for creating a voice generation model based on the voice data, and means for analyzing the user's emotional state and acquiring emotional data. This makes it possible to generate and provide voice data that corresponds to the user's emotions.

[0805] A "voice provider" is a person who has the right to provide voice data and whose data is used to construct a voice generation model.

[0806] A "licensing agreement" is a contract that forms an agreement with the audio provider for the use of specific rights.

[0807] A "speech generation model" is an algorithm or program that generates speech that corresponds to the user's emotions based on the speech data of a speech provider.

[0808] An "emotion analysis device" is a device that analyzes a user's emotional state based on information provided by the user and data from terminal sensors.

[0809] "Voice data generation" is the process of generating voice data to be provided to the user using a voice generation model based on acquired emotional data.

[0810] "Official authentication information" refers to identification information assigned to prove that the generated audio data is legitimate.

[0811] "Usage fees" refer to the fees that users must pay to use the generated audio data.

[0812] "License fees" refer to the payment made to the voice provider for the right to use the voice data and build the model.

[0813] To implement this invention, a voice provider provides voice data to a server, and the server constructs a voice generation model based on that data. A licensing agreement is concluded between the voice provider and the server, securing the right to use the voice data. The server receives emotion data transmitted from the user terminal via an emotion analysis device. This emotion data is acquired through user input and sensors and is used to analyze the user's emotional state in real time.

[0814] The server analyzes the user's emotional state and uses a speech generation model to generate audio data corresponding to that emotional data. The generated audio data is accompanied by official authentication information to guarantee its authenticity. The speech generation model is developed using Python and related libraries (PyTorch, TensorFlow). The generated audio data is transmitted to the user's terminal via the internet.

[0815] The user confirms that the audio data has been provided on their device and pays the usage fee through a digital payment system. The server distributes the usage fee to the audio provider as a licensing fee. This allows users to use audio content that fits their emotions, and audio providers to receive compensation for the audio data they provide.

[0816] For example, when a user is listening to a horror novel in an audiobook application, if the emotion analyzer recognizes the user's level of tension, the voice generated by the model will be adjusted to a gentler tone. This mechanism improves the accuracy of emotion-based voice generation and enhances the user experience. The voice generation process can be controlled using a prompt such as, "Generate voice reading the next text in a manner appropriate for when the user's stress level is high."

[0817] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0818] Step 1:

[0819] The user terminal requests audio content and collects data related to the user's emotions during the process. It receives user voice input and sensor information as input, processes it with an emotion analysis device, and outputs emotion data. Specifically, it uses the user terminal's microphone and camera to capture the user's facial expressions and voice tone in real time.

[0820] Step 2:

[0821] The server receives emotion data sent from the user's terminal. Based on the input emotion data, it uses an emotion analysis engine to analyze the user's emotional state and outputs the results in a data format. Specifically, it analyzes the obtained emotion information and evaluates the user's emotional state (e.g., joy, sadness, tension).

[0822] Step 3:

[0823] The server constructs a speech generation model based on the analyzed emotion data and the user's request text. The emotion data and text are input to the speech generation model, and the adjusted speech data is output. Specifically, the generation AI model built on the server generates a speech waveform corresponding to the emotion.

[0824] Step 4:

[0825] The server assigns official authentication information to the generated audio data. It labels the audio data with authentication information in text format and outputs the result. Specifically, it adds a digital signature to the generated audio file to verify its authenticity.

[0826] Step 5:

[0827] The server sends audio data with official authentication information attached to the user's terminal. The audio data is transferred using a communication protocol, and the audio file is output to the user's terminal. Specifically, the audio data is streamed in real time over the internet.

[0828] Step 6:

[0829] The user receives and confirms the voice data on their device and pays the usage fee. They verify that the voice is as expected as input and pay the fee as output through a digital payment system. Specifically, they confirm the billing information on the application and complete the payment process.

[0830] Step 7:

[0831] The server distributes the collected royalties to the voice providers as licensing fees. Based on the royalty data received as input, it generates payment orders to the voice providers as output. Specifically, it uses an electronic payment platform to automatically process transfers to the voice providers.

[0832] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0833] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0834] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0835] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0836] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0837] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0838] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0839] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0840] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0841] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0842] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0843] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0844] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0845] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0846] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0847] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0848] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0849] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0850] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0851] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0852] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0853] The following is further disclosed regarding the embodiments described above.

[0854] (Claim 1)

[0855] A means of concluding a license agreement with an audio provider to permit the provision of audio data,

[0856] A means for creating a speech generation model based on the voice data of a voice provider,

[0857] A means for generating voice data using the voice generation model in response to a request from a user terminal,

[0858] A means of attaching official authentication information to the generated audio data,

[0859] A method for collecting usage fees from users and distributing them to voice providers as license fees,

[0860] A system that includes this.

[0861] (Claim 2)

[0862] The system according to claim 1, comprising means for providing a user terminal with an interface for generating voice data.

[0863] (Claim 3)

[0864] The system according to claim 1, comprising means for transmitting generated audio data to a user terminal.

[0865] "Example 1"

[0866] (Claim 1)

[0867] A means of concluding a rights agreement with the speaker to permit the provision of sound information,

[0868] A means for creating a speech generation algorithm based on the acoustic information of the speaker,

[0869] A means for generating acoustic information using the voice generation algorithm in response to a request from a user device,

[0870] A means of attaching official authentication information to the generated acoustic information,

[0871] A method of collecting usage fees from users and distributing them to the voice actors as royalties,

[0872] A means for inputting a presentation sentence for generating acoustic information into a speech generation algorithm,

[0873] A means of guaranteeing reliability by attaching official authentication information to acoustic information,

[0874] A means for transmitting acoustic information generated in real time to a user device,

[0875] A system that includes this.

[0876] (Claim 2)

[0877] The system according to claim 1, comprising means for providing a user device with an input / output mechanism for generating acoustic information.

[0878] (Claim 3)

[0879] The system according to claim 1, comprising means for transmitting generated acoustic information to a user device.

[0880] "Application Example 1"

[0881] (Claim 1)

[0882] A means of concluding a license agreement with an audio provider to permit the provision of audio data,

[0883] A means for creating a speech generation model based on the voice data of a voice provider,

[0884] A means for generating voice data using the voice generation model in response to a request from a user terminal,

[0885] A means of attaching official authentication information to the generated acoustic information,

[0886] A method for collecting usage fees from users and distributing them to voice providers as license fees,

[0887] Means for providing an information input area for generating acoustic information,

[0888] A means for transmitting the generated acoustic information to an information terminal,

[0889] A system that includes this.

[0890] (Claim 2)

[0891] The system according to claim 1, comprising means for selecting an acoustic generation model based on information entered by a user into an information input area.

[0892] (Claim 3)

[0893] The system according to claim 1, comprising means for enabling the reception of acoustic information generated in an information terminal.

[0894] "Example 2 of combining an emotion engine"

[0895] (Claim 1)

[0896] A means of concluding a license agreement with an audio provider to permit the provision of audio data,

[0897] A means for creating a speech generation model based on the voice data of a voice provider,

[0898] A means of analyzing the user's emotional state based on input data from the terminal,

[0899] A means for generating voice data using the voice generation model in response to a request from the user terminal and the analyzed emotional state,

[0900] A means of attaching official authentication information to the generated audio data,

[0901] A method for collecting usage fees from users and distributing them to voice providers as license fees,

[0902] A system that includes this.

[0903] (Claim 2)

[0904] The system according to claim 1, comprising means for providing a user terminal with an operation screen for generating voice data.

[0905] (Claim 3)

[0906] The system according to claim 1, comprising means for transmitting generated audio data to a user terminal.

[0907] "Application example 2 of combining emotional engines"

[0908] (Claim 1)

[0909] A means of concluding a licensing agreement with an audio provider to permit the provision of audio data,

[0910] A means for creating a speech generation model based on the voice data of a voice provider,

[0911] A means for acquiring user emotional data using an emotion analysis device that analyzes the user's emotional state,

[0912] A means for generating speech data using a speech generation model in accordance with emotional data based on a request from a user terminal,

[0913] A means of attaching official authentication information to the generated audio data,

[0914] A method of collecting usage fees from users and distributing them to audio providers as licensing fees,

[0915] Information processing device including

[0916] (Claim 2)

[0917] The information processing apparatus according to claim 1, comprising means for providing a user interface for generating voice data to a user terminal.

[0918] (Claim 3)

[0919] The information processing apparatus according to claim 1, comprising means for transmitting generated audio data to a user terminal. [Explanation of Symbols]

[0920] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of concluding a license agreement with an audio provider to permit the provision of audio data, A means for creating a speech generation model based on the voice data of a voice provider, A means for generating voice data using the voice generation model in response to a request from a user terminal, A means of attaching official authentication information to the generated audio data, A method for collecting usage fees from users and distributing them to voice providers as license fees, A system that includes this.

2. The system according to claim 1, comprising means for providing a user terminal with an interface for generating voice data.

3. The system according to claim 1, further comprising means for transmitting generated audio data to a user terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A