System
The system addresses the challenge of generating high-quality audio for advertising by using a learning and generation AI to create voice clones, improving efficiency and reducing costs in the production process.
Patent Information
- Application Number
- JP2024136559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Conventional methods face challenges in quickly generating high-quality audio for advertising production due to schedule and budget constraints.
A system comprising a learning unit, generation unit, and advertisement creation unit that learns voice data, generates high-quality voice clones using a generation AI, and creates advertisements, thereby streamlining the production process.
The system efficiently generates high-quality voice clones, reducing production costs and time, and enhances advertising creativity by eliminating the need for scheduling voice actors and recording studios.
Smart Images

Figure 2026033513000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With conventional technology, it has been difficult to quickly generate high-quality audio for advertising production, and there have been challenges such as schedule and budget constraints.
[0005] The system according to the embodiment aims to quickly generate high-quality voice clones and realize efficient advertisement production. [Means for solving the problem]
[0006] A system according to an embodiment includes a learning unit, a generation unit, and an advertisement creation unit. The learning unit learns voice data. The generation unit generates voice clones based on the data learned by the learning unit. The advertisement creation unit creates advertisements using the voice clones generated by the generation unit. [Effects of the Invention]
[0007] The system according to the embodiment can quickly generate high-quality voice clones and realize efficient advertisement production. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION
[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0010] First, the terms used in the following description will be explained.
[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).
[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.
[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example 1) A voice clone generation system according to an embodiment of the present invention utilizes a generation AI to learn specific voices of voice actors or characters and quickly generate high-quality voice clones. This system realizes a flexible and efficient advertising production process and enables reproduction of original voices. For example, the generation AI learns voice data of a voice actor or character and identifies their characteristics. Then, based on the learned data, the generation AI generates a high-quality voice clone. This voice clone is used as the original voice in advertising production. This overcomes schedule and budget challenges and improves advertising creativity. The voice clone generation system thus quickly generates high-quality voice clones and improves advertising production efficiency. For example, it eliminates the hassle of scheduling voice actors and arranging recording studios, thereby reducing advertising production costs. Furthermore, the rapid generation of high-quality voice clones shortens advertising production schedules and enables timely advertising deployment.
[0029] A voice clone generation system according to an embodiment includes a learning unit, a generation unit, and an advertisement production unit. The learning unit learns voice data of voice actors and characters. For example, the learning unit analyzes in detail the tone, rhythm, pronunciation characteristics, etc. of the voice to understand these characteristics. The generation unit generates a voice clone based on the data learned by the learning unit. For example, the generation unit generates a high-quality voice clone using a generation AI. The generated voice clone has a quality that is almost indistinguishable from the original voice. The advertisement production unit produces advertisements using the voice clone generated by the generation unit. For example, the advertisement production unit uses the generated voice clone to produce advertisement narration and character lines. This makes it possible to streamline the processes of learning voice data, generating voice clones, and producing advertisements.
[0030] The learning unit can analyze the voice data of a voice actor or character using a specific method to grasp its characteristics. The learning unit, for example, analyzes the voice data using spectral analysis. For example, the learning unit analyzes the frequency components of the voice data to grasp the tone and rhythm of the voice. The learning unit can also analyze the voice data using acoustic feature extraction. For example, the learning unit extracts the pitch and formants of the voice data to grasp the pronunciation characteristics. The learning unit can also analyze the voice data using deep learning. For example, the learning unit inputs the voice data into a neural network to learn the voice characteristics. This improves the accuracy of the voice clone through detailed analysis of the voice data.
[0031] The generation unit can generate highly accurate voice clones based on the learned data. The generation unit generates voice clones using, for example, a generation AI. For example, the generation unit generates high-quality voice clones using voice synthesis technology. The generation unit can also evaluate the quality of the voice clones based on the voice recognition rate. For example, the generation unit measures the recognition rate of the generated voice clones and evaluates the quality. The generation unit can also evaluate the quality of the voice clones based on sound quality evaluation. For example, the generation unit evaluates the sound quality of the generated voice clones and improves the quality. This generates high-quality voice clones, thereby improving the quality of the advertisement.
[0032] The advertisement production department can use the generated voice clone to create a narration or character lines for the advertisement. For example, the advertisement production department creates a narration for the advertisement using the generated voice clone. For example, the advertisement production department uses the voice clone based on the content of the script to record a narration. The advertisement production department can also create character lines using the generated voice clone. For example, the advertisement production department uses the voice clone based on the character's settings to record lines. The advertisement production department can also edit the audio for the advertisement using the generated voice clone. For example, the advertisement production department edits the audio for the advertisement to optimize it. As a result, the efficiency of advertisement production is improved by using the generated voice clone.
[0033] The generation unit can generate a voice clone in a short time. For example, the generation unit generates a voice clone in a few seconds. For example, the generation unit uses a high-speed generation AI to quickly generate a voice clone. The generation unit can also generate a voice clone in a few minutes. For example, the generation unit uses an efficient algorithm to quickly generate a voice clone. The generation unit can also generate a voice clone in real time. For example, the generation unit uses real-time processing technology to instantly generate a voice clone. This allows for the rapid generation of voice clones, thereby shortening the schedule for advertising production.
[0034] The advertising production department can shorten the advertisement production schedule by using the generated voice clone. For example, the advertising production department shortens the advertisement production schedule by using the generated voice clone. For example, the advertising production department quickly generates voice clones to streamline each process of advertisement production. The advertising production department can also optimize advertisement production resources by using the generated voice clone. For example, the advertising production department can use the voice clone to eliminate the need to arrange a recording studio and shorten the schedule. The advertising production department can also shorten the delivery time for advertisement production by using the generated voice clone. For example, the advertising production department can use the voice clone to quickly complete an advertisement and meet the delivery date. This shortens the advertisement production schedule, thereby enabling timely advertisement deployment.
[0035] When learning voice data, the learning unit can analyze the past acting history of the voice actor or character and select an appropriate learning method. The learning unit, for example, analyzes the past acting history of the voice actor. For example, the learning unit analyzes past voice data and learns pronunciation characteristics in specific scenes. The learning unit can also analyze the past acting history of the character. For example, the learning unit analyzes past acting data and learns patterns of emotional expression. The learning unit can also learn specific tones and rhythms. For example, the learning unit analyzes past acting data and learns specific tones and rhythms. In this way, by analyzing the past acting history, the optimal learning method can be selected and the accuracy of learning can be improved.
[0036] When learning the voice data, the learning unit can perform filtering based on specific scenes or situations of the voice actors or characters. The learning unit, for example, filters voice data in specific scenes. For example, the learning unit extracts voice data in specific scenes to narrow down the learning target. The learning unit can also filter voice data in specific situations. For example, the learning unit extracts voice data in specific situations to improve the accuracy of learning. The learning unit can also filter voice data based on specific emotional expressions. For example, the learning unit extracts voice data based on specific emotional expressions to improve the efficiency of learning. In this way, filtering based on specific scenes or situations can improve the accuracy of learning.
[0037] When learning the audio data, the learning unit can perform analysis using a specific method for capturing subtle changes in the pronunciation of a voice actor or character. The learning unit, for example, analyzes subtle changes in the voice actor's pronunciation in detail. For example, the learning unit analyzes changes in audio waveforms to capture subtle changes in pronunciation. The learning unit can also analyze fluctuations in acoustic features. For example, the learning unit analyzes fluctuations in acoustic features to capture subtle changes in pronunciation. The learning unit can also analyze subtle changes in pronunciation in detail in a specific scene. For example, the learning unit analyzes audio data in a specific scene to capture subtle changes in pronunciation. This allows for detailed analysis of subtle changes in pronunciation, thereby improving the accuracy of learning.
[0038] When learning voice data, the learning unit can prioritize learning highly relevant data by taking into account geographical background information of the voice actor or character. The learning unit, for example, takes into account geographical background information of the voice actor. For example, the learning unit prioritizes learning highly relevant voice data based on the voice actor's hometown or area of activity. The learning unit can also consider geographical background information of the character. For example, the learning unit prioritizes learning highly relevant voice data based on the character's settings. The learning unit can also prioritize learning voice data from a specific region. For example, the learning unit learns voice data by taking into account the cultural and linguistic characteristics of a specific region. In this way, by taking geographical background information into account, highly relevant data can be prioritized and the accuracy of learning can be improved.
[0039] When learning the voice data, the learning unit can analyze the social media activities of the voice actor or character and learn related data. The learning unit, for example, analyzes the social media activities of the voice actor. For example, the learning unit analyzes the content of the voice actor's posts and the reactions of followers to learn related voice data. The learning unit can also analyze the social media activities of the character. For example, the learning unit analyzes the content of the character's social media activities and learns related voice data. The learning unit can also learn voice data based on specific social media activities. For example, the learning unit learns voice data related to specific events or campaigns. In this way, by analyzing social media activities, related data can be learned and the accuracy of learning can be improved.
[0040] When learning voice data, the learning unit can customize the learning method by reflecting past feedback on the voice actor or character. The learning unit, for example, reflects past feedback on the voice actor. For example, the learning unit customizes the learning method based on user ratings and comments. The learning unit can also reflect past feedback on the character. For example, the learning unit customizes the learning method based on feedback on the character's performance. The learning unit can also customize the learning method based on specific feedback. For example, the learning unit adjusts the learning method based on feedback on a specific scene or situation. In this way, by reflecting past feedback, the learning method can be customized and the accuracy of learning can be improved.
[0041] When generating a voice clone, the generation unit can adjust the level of detail of the voice clone based on important features of the voice actor or character. For example, the generation unit adjusts the level of detail of the voice clone based on important features of the voice actor. For example, the generation unit adjusts the level of detail of the voice clone based on features such as voice pitch, rhythm, and intonation. The generation unit can also adjust the level of detail of the voice clone based on important features of the character. For example, the generation unit adjusts the level of detail of the voice clone based on the character's setting and personality. The generation unit can also adjust the level of detail of the voice clone based on important features in a particular scene. For example, the generation unit adjusts the level of detail of the voice clone based on the voice tone and emotional expression in a particular scene. In this way, adjusting the level of detail of the voice clone based on important features improves the quality of the voice clone.
[0042] When generating a voice clone, the generation unit can apply different generation algorithms depending on the voice actor or character category. For example, the generation unit applies different generation algorithms depending on the voice actor category. For example, the generation unit selects an optimal generation algorithm depending on the voice actor's genre or role. The generation unit can also apply different generation algorithms depending on the character category. For example, the generation unit selects an optimal generation algorithm depending on the character type or setting. The generation unit can also apply different generation algorithms depending on the category of a specific scene. For example, the generation unit selects an optimal generation algorithm depending on the emotional expression or situation in a specific scene. In this way, applying different generation algorithms depending on the category improves the quality of the voice clone.
[0043] When generating a voice clone, the generation unit can improve the accuracy of generation by referring to past generation results of a voice actor or character. The generation unit, for example, refers to past generation results of a voice actor. For example, the generation unit analyzes past voice clone data to improve the accuracy of generation. The generation unit can also refer to past generation results of a character. For example, the generation unit analyzes past generation data to improve the accuracy of generation. The generation unit can also refer to past generation results of a specific scene. For example, the generation unit analyzes past generation data of a specific scene to improve the accuracy of generation. In this way, the accuracy of the voice clone is improved by referring to past generation results.
[0044] When generating voice clones, the generation unit can determine the generation priority based on the performance period of the voice actor or character. The generation unit determines the generation priority, for example, based on the performance period of the voice actor. For example, the generation unit determines the generation priority based on performance data for a specific year or season. The generation unit can also determine the generation priority based on the performance period of the character. For example, the generation unit determines the generation priority based on performance data for a specific season or event. The generation unit can also determine the generation priority based on the performance period of a specific scene. For example, the generation unit determines the generation priority based on performance data for a specific scene. In this way, by determining the generation priority based on the performance period, efficient voice clone generation is possible.
[0045] When generating voice clones, the generation unit can adjust the order of generation based on the relevance of voice actors or characters. The generation unit adjusts the order of generation based on, for example, the relevance of voice actors. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story. The generation unit can also adjust the order of generation based on the relevance of characters. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story. The generation unit can also adjust the order of generation based on the relevance in a specific scene. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story in a specific scene. This enables efficient voice clone generation by adjusting the order of generation based on the relevance.
[0046] When generating a voice clone, the generation unit can adjust the use of technical terminology in the generation according to the expertise level of the voice actor or character. The generation unit, for example, adjusts the use of technical terminology in the voice clone according to the expertise level of the voice actor. For example, the generation unit uses simple words for beginners and detailed technical terminology for advanced users. The generation unit can also adjust the use of technical terminology in the voice clone according to the expertise level of the character. For example, the generation unit adjusts the use of technical terminology based on the character's settings. The generation unit can also adjust the use of technical terminology in the voice clone according to the expertise level in a particular scene. For example, the generation unit adjusts the use of technical terminology based on the expertise level in a particular scene. In this way, by adjusting the use of technical terminology according to the expertise level, a more appropriate voice clone can be generated.
[0047] The advertisement production unit can adjust the level of detail of production based on the quality of the generated voice clone when producing an advertisement. The advertisement production unit adjusts the level of detail of production based on, for example, the quality of the generated voice clone. For example, the advertisement production unit produces a detailed advertisement based on a high-quality voice clone. The advertisement production unit can also produce an advertisement with a moderate level of detail based on a medium-quality voice clone. The advertisement production unit can also produce a simplified advertisement based on a low-quality voice clone. In this way, adjusting the level of detail of production based on the quality of the voice clone improves the quality of the advertisement.
[0048] The advertising production department can apply different production techniques depending on the category of the advertisement when creating the advertisement. For example, the advertising production department applies a visually appealing technique to an advertisement for a product. For example, the advertising production department uses visual effects to emphasize the features of the product. The advertising production department can also apply an explanatory technique to an advertisement for a service. For example, the advertising production department uses detailed narration to explain the benefits of the service. The advertising production department can also apply an emotionally appealing technique to an advertisement for an event. For example, the advertising production department uses emotional music and images to convey the atmosphere of the event. In this way, by applying different production techniques depending on the category of the advertisement, more effective advertisements can be created.
[0049] When creating an advertisement, the advertising production department can improve the accuracy of production by referring to past advertising production results. The advertising production department, for example, refers to past successful advertising production results. For example, the advertising production department analyzes the evaluations and viewer responses of past advertisements and creates new advertisements. The advertising production department can also analyze past unsuccessful advertising production results and create new advertisements by reflecting the improvements. For example, the advertising production department identifies problems with past advertisements and adopts methods to improve them. The advertising production department can also select the optimal production method based on past advertising production results and create new advertisements. For example, the advertising production department analyzes the factors that made past advertisements successful and reflects them in new advertisements. In this way, the accuracy of advertisements can be improved by referring to past advertising production results.
[0050] The advertisement production department can determine the priority of production based on the submission time of the generated voice clones when producing an advertisement. The advertisement production department determines the priority of production based on, for example, the submission time of the generated voice clones. For example, the advertisement production department produces an advertisement by preferentially using voice clones submitted early. The advertisement production department can also adjust the schedule of advertisement production based on the submission time. For example, the advertisement production department optimally allocates advertisement production resources according to the submission time. The advertisement production department can also adjust the delivery date of advertisement production based on the submission time. For example, the advertisement production department streamlines each process of advertisement production according to the submission time. As a result, by determining the priority of production based on the submission time, efficient advertisement production is possible.
[0051] The advertisement production unit can adjust the order of production based on the relevance of the generated voice clones when producing an advertisement. The advertisement production unit adjusts the order of production based on, for example, the relevance of the generated voice clones. For example, the advertisement production unit produces an advertisement by preferentially using highly relevant voice clones. The advertisement production unit can also adjust the order of advertisement production based on the relevance of the voice clones. For example, the advertisement production unit produces an advertisement by putting less relevant voice clones on hold. The advertisement production unit can also adjust the order of production based on the relevance in a specific scene. For example, the advertisement production unit adjusts the order of production based on the relevance of the voice clones in a specific scene. As a result, adjusting the order of production based on the relevance enables efficient advertisement production.
[0052] The advertisement production unit can adjust the use of technical terms in the production according to the user's level of expertise. For example, the advertisement production unit adjusts the use of technical terms in the production according to the user's level of expertise. For example, the advertisement production unit uses simple words for beginners and detailed technical terms for advanced users. The advertisement production unit can also adjust the content of the advertisement according to the user's level of expertise. For example, the advertisement production unit provides detailed technical information for users with high expertise and basic information for users with low expertise. The advertisement production unit can also adjust the content of the advertisement according to the level of expertise in a particular scene. For example, the advertisement production unit adjusts the content of the advertisement based on the user's level of expertise in a particular scene. In this way, by adjusting the use of technical terms according to the level of expertise, a more appropriate advertisement can be produced.
[0053] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0054] When learning voice data, the learning unit can also take into account the health status of the voice actor or character. For example, if the voice actor has a cold, that voice data can be excluded during learning. Also, if the voice actor is tired, that voice data can be filtered during learning. Furthermore, it can prioritize learning of voice data recorded by the voice actor in optimal condition. This allows for learning that takes into account the health status of the voice, making it possible to generate higher quality voice clones.
[0055] The advertising production department can also use the generated voice clones to customize the advertisement according to the target audience. For example, for an advertisement targeted at younger generations, the generated voice clone can be used to provide narration in a casual tone. For an advertisement targeted at older people, the generated voice clone can be used to provide narration in a more subdued tone. Furthermore, for advertisements targeted at specific regions, voice clones incorporating the dialect or accent of that region can be used. This allows for customization according to the target audience, maximizing the effectiveness of the advertisement.
[0056] The advertising production department can use the generated voice clones to add sound effects according to the ad scenario. For example, in action scenes, echo and reverb can be added to the generated voice clones to enhance realism. In emotional scenes, soft effects can be added to the generated voice clones to enhance emotions. Furthermore, in comedy scenes, comical effects can be added to the generated voice clones to elicit laughter. In this way, adding sound effects according to the scenario can enhance the appeal of the advertisement.
[0057] The generation unit can also adjust the texture of the voice of the voice actor or character when generating a voice clone. For example, if the voice actor has a soft voice, the sound quality of the voice clone can be adjusted to reproduce that texture. Also, if the voice actor has a husky voice, the sound quality of the voice clone can be adjusted to reproduce that texture. Furthermore, if the voice actor has a clear voice, the sound quality of the voice clone can be adjusted to reproduce that texture. In this way, by adjusting the texture of the voice, a more realistic voice clone can be generated.
[0058] The advertising production department can use the generated voice clones to add music and sound effects according to the theme of the advertisement. For example, for an emotional advertisement, emotional music can be added to the generated voice clones. For an action advertisement, powerful sound effects can be added to the generated voice clones. Furthermore, for a comedy advertisement, comical music and sound effects can be added to the generated voice clones. In this way, the appeal of the advertisement can be enhanced by adding music and sound effects according to the theme.
[0059] The processing flow of the first embodiment will be briefly explained below.
[0060] Step 1: The learning unit studies the voice data of voice actors and characters. For example, the learning unit analyzes in detail the tone, rhythm, and pronunciation characteristics of the voice to understand their characteristics. Step 2: The generator generates a voice clone based on the data learned by the trainer. For example, the generator uses a generation AI to generate a high-quality voice clone. The generated voice clone has a quality that is almost indistinguishable from the original voice. Step 3: The advertising production department uses the voice clone generated by the generation department to create an advertisement. For example, the advertising production department uses the generated voice clone to create the narration and character lines of the advertisement. This makes it possible to learn voice data, generate voice clones, and streamline the process of creating advertisements.
[0061] (Example 2) A voice clone generation system according to an embodiment of the present invention utilizes a generation AI to learn specific voices of voice actors or characters and quickly generate high-quality voice clones. This system realizes a flexible and efficient advertising production process and enables reproduction of original voices. For example, the generation AI learns voice data of a voice actor or character and identifies their characteristics. Then, based on the learned data, the generation AI generates a high-quality voice clone. This voice clone is used as the original voice in advertising production. This overcomes schedule and budget challenges and improves advertising creativity. The voice clone generation system thus quickly generates high-quality voice clones and improves advertising production efficiency. For example, it eliminates the hassle of scheduling voice actors and arranging recording studios, thereby reducing advertising production costs. Furthermore, the rapid generation of high-quality voice clones shortens advertising production schedules and enables timely advertising deployment.
[0062] A voice clone generation system according to an embodiment includes a learning unit, a generation unit, and an advertisement production unit. The learning unit learns voice data of voice actors and characters. For example, the learning unit analyzes in detail the tone, rhythm, pronunciation characteristics, etc. of the voice to understand these characteristics. The generation unit generates a voice clone based on the data learned by the learning unit. For example, the generation unit generates a high-quality voice clone using a generation AI. The generated voice clone has a quality that is almost indistinguishable from the original voice. The advertisement production unit produces advertisements using the voice clone generated by the generation unit. For example, the advertisement production unit uses the generated voice clone to produce advertisement narration and character lines. This makes it possible to streamline the processes of learning voice data, generating voice clones, and producing advertisements.
[0063] The learning unit can analyze the voice data of a voice actor or character using a specific method to grasp its characteristics. The learning unit, for example, analyzes the voice data using spectral analysis. For example, the learning unit analyzes the frequency components of the voice data to grasp the tone and rhythm of the voice. The learning unit can also analyze the voice data using acoustic feature extraction. For example, the learning unit extracts the pitch and formants of the voice data to grasp the pronunciation characteristics. The learning unit can also analyze the voice data using deep learning. For example, the learning unit inputs the voice data into a neural network to learn the voice characteristics. This improves the accuracy of the voice clone through detailed analysis of the voice data.
[0064] The generation unit can generate highly accurate voice clones based on the learned data. The generation unit generates voice clones using, for example, a generation AI. For example, the generation unit generates high-quality voice clones using voice synthesis technology. The generation unit can also evaluate the quality of the voice clones based on the voice recognition rate. For example, the generation unit measures the recognition rate of the generated voice clones and evaluates the quality. The generation unit can also evaluate the quality of the voice clones based on sound quality evaluation. For example, the generation unit evaluates the sound quality of the generated voice clones and improves the quality. This generates high-quality voice clones, thereby improving the quality of the advertisement.
[0065] The advertisement production department can use the generated voice clone to create a narration or character lines for the advertisement. For example, the advertisement production department creates a narration for the advertisement using the generated voice clone. For example, the advertisement production department uses the voice clone based on the content of the script to record a narration. The advertisement production department can also create character lines using the generated voice clone. For example, the advertisement production department uses the voice clone based on the character's settings to record lines. The advertisement production department can also edit the audio for the advertisement using the generated voice clone. For example, the advertisement production department edits the audio for the advertisement to optimize it. As a result, the efficiency of advertisement production is improved by using the generated voice clone.
[0066] The generation unit can generate a voice clone in a short time. For example, the generation unit generates a voice clone in a few seconds. For example, the generation unit uses a high-speed generation AI to quickly generate a voice clone. The generation unit can also generate a voice clone in a few minutes. For example, the generation unit uses an efficient algorithm to quickly generate a voice clone. The generation unit can also generate a voice clone in real time. For example, the generation unit uses real-time processing technology to instantly generate a voice clone. This allows for the rapid generation of voice clones, thereby shortening the schedule for advertising production.
[0067] The advertising production department can shorten the advertisement production schedule by using the generated voice clone. For example, the advertising production department shortens the advertisement production schedule by using the generated voice clone. For example, the advertising production department quickly generates voice clones to streamline each process of advertisement production. The advertising production department can also optimize advertisement production resources by using the generated voice clone. For example, the advertising production department can use the voice clone to eliminate the need to arrange a recording studio and shorten the schedule. The advertising production department can also shorten the delivery time for advertisement production by using the generated voice clone. For example, the advertising production department can use the voice clone to quickly complete an advertisement and meet the delivery date. This shortens the advertisement production schedule, thereby enabling timely advertisement deployment.
[0068] The learning unit can estimate the user's emotions and adjust the timing of learning the voice data based on the estimated user's emotions. The learning unit, for example, estimates the user's emotions. For example, the learning unit estimates the user's emotions using an emotion recognition algorithm. The learning unit also adjusts the timing of learning the voice data based on the estimated user's emotions. For example, the learning unit can cause the generation AI to learn the voice data at night when the user is relaxed. The learning unit can also cause the generation AI to learn the voice data in a short period of time when the user is feeling stressed. The learning unit can also cause the generation AI to learn the voice data continuously when the user is concentrating. This enables efficient learning by adjusting the learning timing according to the user's emotions.
[0069] When learning voice data, the learning unit can analyze the past acting history of the voice actor or character and select an appropriate learning method. The learning unit, for example, analyzes the past acting history of the voice actor. For example, the learning unit analyzes past voice data and learns pronunciation characteristics in specific scenes. The learning unit can also analyze the past acting history of the character. For example, the learning unit analyzes past acting data and learns patterns of emotional expression. The learning unit can also learn specific tones and rhythms. For example, the learning unit analyzes past acting data and learns specific tones and rhythms. In this way, by analyzing the past acting history, the optimal learning method can be selected and the accuracy of learning can be improved.
[0070] When learning the voice data, the learning unit can perform filtering based on specific scenes or situations of the voice actors or characters. The learning unit, for example, filters voice data in specific scenes. For example, the learning unit extracts voice data in specific scenes to narrow down the learning target. The learning unit can also filter voice data in specific situations. For example, the learning unit extracts voice data in specific situations to improve the accuracy of learning. The learning unit can also filter voice data based on specific emotional expressions. For example, the learning unit extracts voice data based on specific emotional expressions to improve the efficiency of learning. In this way, filtering based on specific scenes or situations can improve the accuracy of learning.
[0071] When learning the audio data, the learning unit can perform analysis using a specific method for capturing subtle changes in the pronunciation of a voice actor or character. The learning unit, for example, analyzes subtle changes in the voice actor's pronunciation in detail. For example, the learning unit analyzes changes in audio waveforms to capture subtle changes in pronunciation. The learning unit can also analyze fluctuations in acoustic features. For example, the learning unit analyzes fluctuations in acoustic features to capture subtle changes in pronunciation. The learning unit can also analyze subtle changes in pronunciation in detail in a specific scene. For example, the learning unit analyzes audio data in a specific scene to capture subtle changes in pronunciation. This allows for detailed analysis of subtle changes in pronunciation, thereby improving the accuracy of learning.
[0072] The learning unit can estimate the user's emotions and determine the priority of the voice data to be learned based on the estimated user's emotions. The learning unit, for example, estimates the user's emotions. For example, the learning unit estimates the user's emotions using an emotion recognition algorithm. The learning unit also determines the priority of the voice data to be learned based on the estimated user's emotions. For example, if the user is relaxed, the learning unit can prioritize learning voice data that makes the generation AI relaxed. Also, if the user is feeling stressed, the learning unit can also prioritize learning voice data that reduces stress. Also, if the user is concentrating, the learning unit can prioritize learning voice data that increases concentration. In this way, efficient learning is possible by determining the priority of the voice data to be learned according to the user's emotions.
[0073] When learning voice data, the learning unit can prioritize learning highly relevant data by taking into account geographical background information of the voice actor or character. The learning unit, for example, takes into account geographical background information of the voice actor. For example, the learning unit prioritizes learning highly relevant voice data based on the voice actor's hometown or area of activity. The learning unit can also consider geographical background information of the character. For example, the learning unit prioritizes learning highly relevant voice data based on the character's settings. The learning unit can also prioritize learning voice data from a specific region. For example, the learning unit learns voice data by taking into account the cultural and linguistic characteristics of a specific region. In this way, by taking geographical background information into account, highly relevant data can be prioritized and the accuracy of learning can be improved.
[0074] When learning the voice data, the learning unit can analyze the social media activities of the voice actor or character and learn related data. The learning unit, for example, analyzes the social media activities of the voice actor. For example, the learning unit analyzes the content of the voice actor's posts and the reactions of followers to learn related voice data. The learning unit can also analyze the social media activities of the character. For example, the learning unit analyzes the content of the character's social media activities and learns related voice data. The learning unit can also learn voice data based on specific social media activities. For example, the learning unit learns voice data related to specific events or campaigns. In this way, by analyzing social media activities, related data can be learned and the accuracy of learning can be improved.
[0075] When learning voice data, the learning unit can customize the learning method by reflecting past feedback on the voice actor or character. The learning unit, for example, reflects past feedback on the voice actor. For example, the learning unit customizes the learning method based on user ratings and comments. The learning unit can also reflect past feedback on the character. For example, the learning unit customizes the learning method based on feedback on the character's performance. The learning unit can also customize the learning method based on specific feedback. For example, the learning unit adjusts the learning method based on feedback on a specific scene or situation. In this way, by reflecting past feedback, the learning method can be customized and the accuracy of learning can be improved.
[0076] The generation unit can estimate the user's emotion and adjust the expression method of the generated voice clone based on the estimated user's emotion. The generation unit, for example, estimates the user's emotion. For example, the generation unit estimates the user's emotion using an emotion recognition algorithm. The generation unit also adjusts the expression method of the voice clone based on the estimated user's emotion. For example, if the user is relaxed, the generation unit causes the generation AI to generate a voice clone using a relaxed expression method. Also, if the user is feeling stressed, the generation unit can generate a voice clone using an expression method that reduces stress. Also, if the user is concentrating, the generation unit can generate a voice clone using an expression method that increases concentration. In this way, by adjusting the expression method of the voice clone according to the user's emotion, a more appropriate voice clone can be generated.
[0077] When generating a voice clone, the generation unit can adjust the level of detail of the voice clone based on important features of the voice actor or character. For example, the generation unit adjusts the level of detail of the voice clone based on important features of the voice actor. For example, the generation unit adjusts the level of detail of the voice clone based on features such as voice pitch, rhythm, and intonation. The generation unit can also adjust the level of detail of the voice clone based on important features of the character. For example, the generation unit adjusts the level of detail of the voice clone based on the character's setting and personality. The generation unit can also adjust the level of detail of the voice clone based on important features in a particular scene. For example, the generation unit adjusts the level of detail of the voice clone based on the voice tone and emotional expression in a particular scene. In this way, adjusting the level of detail of the voice clone based on important features improves the quality of the voice clone.
[0078] When generating a voice clone, the generation unit can apply different generation algorithms depending on the voice actor or character category. For example, the generation unit applies different generation algorithms depending on the voice actor category. For example, the generation unit selects an optimal generation algorithm depending on the voice actor's genre or role. The generation unit can also apply different generation algorithms depending on the character category. For example, the generation unit selects an optimal generation algorithm depending on the character type or setting. The generation unit can also apply different generation algorithms depending on the category of a specific scene. For example, the generation unit selects an optimal generation algorithm depending on the emotional expression or situation in a specific scene. In this way, applying different generation algorithms depending on the category improves the quality of the voice clone.
[0079] When generating a voice clone, the generation unit can improve the accuracy of generation by referring to past generation results of a voice actor or character. The generation unit, for example, refers to past generation results of a voice actor. For example, the generation unit analyzes past voice clone data to improve the accuracy of generation. The generation unit can also refer to past generation results of a character. For example, the generation unit analyzes past generation data to improve the accuracy of generation. The generation unit can also refer to past generation results of a specific scene. For example, the generation unit analyzes past generation data of a specific scene to improve the accuracy of generation. In this way, the accuracy of the voice clone is improved by referring to past generation results.
[0080] The generation unit can estimate the user's emotion and adjust the length of the generated voice clone based on the estimated user's emotion. The generation unit, for example, estimates the user's emotion. For example, the generation unit estimates the user's emotion using an emotion recognition algorithm. The generation unit also adjusts the length of the voice clone based on the estimated user's emotion. For example, if the user is relaxed, the generation AI can generate a longer voice clone. Also, if the user is feeling stressed, the generation unit can generate a shorter voice clone. Also, if the user is concentrating, the generation AI can generate a voice clone of an appropriate length. In this way, by adjusting the length of the voice clone according to the user's emotion, more appropriate voice clones can be generated.
[0081] When generating voice clones, the generation unit can determine the generation priority based on the performance period of the voice actor or character. The generation unit determines the generation priority, for example, based on the performance period of the voice actor. For example, the generation unit determines the generation priority based on performance data for a specific year or season. The generation unit can also determine the generation priority based on the performance period of the character. For example, the generation unit determines the generation priority based on performance data for a specific season or event. The generation unit can also determine the generation priority based on the performance period of a specific scene. For example, the generation unit determines the generation priority based on performance data for a specific scene. In this way, by determining the generation priority based on the performance period, efficient voice clone generation is possible.
[0082] When generating voice clones, the generation unit can adjust the order of generation based on the relevance of voice actors or characters. The generation unit adjusts the order of generation based on, for example, the relevance of voice actors. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story. The generation unit can also adjust the order of generation based on the relevance of characters. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story. The generation unit can also adjust the order of generation based on the relevance in a specific scene. For example, the generation unit adjusts the order of generation based on the relationship between characters or the relevance of a story in a specific scene. This enables efficient voice clone generation by adjusting the order of generation based on the relevance.
[0083] When generating a voice clone, the generation unit can adjust the use of technical terminology in the generation according to the expertise level of the voice actor or character. The generation unit, for example, adjusts the use of technical terminology in the voice clone according to the expertise level of the voice actor. For example, the generation unit uses simple words for beginners and detailed technical terminology for advanced users. The generation unit can also adjust the use of technical terminology in the voice clone according to the expertise level of the character. For example, the generation unit adjusts the use of technical terminology based on the character's settings. The generation unit can also adjust the use of technical terminology in the voice clone according to the expertise level in a particular scene. For example, the generation unit adjusts the use of technical terminology based on the expertise level in a particular scene. In this way, by adjusting the use of technical terminology according to the expertise level, a more appropriate voice clone can be generated.
[0084] The advertising production department can estimate a user's emotions and adjust the advertisement production method based on the estimated user's emotions. The advertising production department, for example, estimates a user's emotions. For example, the advertising production department estimates a user's emotions using an emotion recognition algorithm. The advertising production department also adjusts the advertisement production method based on the estimated user's emotions. For example, if the user is relaxed, the advertising production department can have the generation AI create a relaxing advertisement. Also, if the user is feeling stressed, the advertising production department can have the generation AI create an advertisement that reduces stress. Also, if the user is concentrating, the advertising production department can have the generation AI create an advertisement that increases concentration. In this way, by adjusting the advertisement production method according to the user's emotions, more effective advertisements can be produced.
[0085] The advertisement production unit can adjust the level of detail of production based on the quality of the generated voice clone when producing an advertisement. The advertisement production unit adjusts the level of detail of production based on, for example, the quality of the generated voice clone. For example, the advertisement production unit produces a detailed advertisement based on a high-quality voice clone. The advertisement production unit can also produce an advertisement with a moderate level of detail based on a medium-quality voice clone. The advertisement production unit can also produce a simplified advertisement based on a low-quality voice clone. In this way, adjusting the level of detail of production based on the quality of the voice clone improves the quality of the advertisement.
[0086] The advertising production department can apply different production techniques depending on the category of the advertisement when creating the advertisement. For example, the advertising production department applies a visually appealing technique to an advertisement for a product. For example, the advertising production department uses visual effects to emphasize the features of the product. The advertising production department can also apply an explanatory technique to an advertisement for a service. For example, the advertising production department uses detailed narration to explain the benefits of the service. The advertising production department can also apply an emotionally appealing technique to an advertisement for an event. For example, the advertising production department uses emotional music and images to convey the atmosphere of the event. In this way, by applying different production techniques depending on the category of the advertisement, more effective advertisements can be created.
[0087] When creating an advertisement, the advertising production department can improve the accuracy of production by referring to past advertising production results. The advertising production department, for example, refers to past successful advertising production results. For example, the advertising production department analyzes the evaluations and viewer responses of past advertisements and creates new advertisements. The advertising production department can also analyze past unsuccessful advertising production results and create new advertisements by reflecting the improvements. For example, the advertising production department identifies problems with past advertisements and adopts methods to improve them. The advertising production department can also select the optimal production method based on past advertising production results and create new advertisements. For example, the advertising production department analyzes the factors that made past advertisements successful and reflects them in new advertisements. In this way, the accuracy of advertisements can be improved by referring to past advertising production results.
[0088] The advertising production department can estimate a user's emotions and adjust the advertisement production schedule based on the estimated user's emotions. The advertising production department, for example, estimates a user's emotions. For example, the advertising production department estimates a user's emotions using an emotion recognition algorithm. The advertising production department also adjusts the advertisement production schedule based on the estimated user's emotions. For example, if the user is relaxed, the advertising production department can have the generation AI create an advertisement with a relaxed schedule. Also, if the user is feeling stressed, the advertising production department can have the generation AI create an advertisement quickly. Also, if the user is concentrating, the advertising production department can have the generation AI create an advertisement with a schedule that makes use of the user's concentration. This enables efficient advertisement production by adjusting the advertisement production schedule according to the user's emotions.
[0089] The advertisement production department can determine the priority of production based on the submission time of the generated voice clones when producing an advertisement. The advertisement production department determines the priority of production based on, for example, the submission time of the generated voice clones. For example, the advertisement production department produces an advertisement by preferentially using voice clones submitted early. The advertisement production department can also adjust the schedule of advertisement production based on the submission time. For example, the advertisement production department optimally allocates advertisement production resources according to the submission time. The advertisement production department can also adjust the delivery date of advertisement production based on the submission time. For example, the advertisement production department streamlines each process of advertisement production according to the submission time. As a result, by determining the priority of production based on the submission time, efficient advertisement production is possible.
[0090] The advertisement production unit can adjust the order of production based on the relevance of the generated voice clones when producing an advertisement. The advertisement production unit adjusts the order of production based on, for example, the relevance of the generated voice clones. For example, the advertisement production unit produces an advertisement by preferentially using highly relevant voice clones. The advertisement production unit can also adjust the order of advertisement production based on the relevance of the voice clones. For example, the advertisement production unit produces an advertisement by putting less relevant voice clones on hold. The advertisement production unit can also adjust the order of production based on the relevance in a specific scene. For example, the advertisement production unit adjusts the order of production based on the relevance of the voice clones in a specific scene. As a result, adjusting the order of production based on the relevance enables efficient advertisement production.
[0091] The advertisement production unit can adjust the use of technical terms in the production according to the user's level of expertise. For example, the advertisement production unit adjusts the use of technical terms in the production according to the user's level of expertise. For example, the advertisement production unit uses simple words for beginners and detailed technical terms for advanced users. The advertisement production unit can also adjust the content of the advertisement according to the user's level of expertise. For example, the advertisement production unit provides detailed technical information for users with high expertise and basic information for users with low expertise. The advertisement production unit can also adjust the content of the advertisement according to the level of expertise in a particular scene. For example, the advertisement production unit adjusts the content of the advertisement based on the user's level of expertise in a particular scene. In this way, by adjusting the use of technical terms according to the level of expertise, a more appropriate advertisement can be produced. === Hard Collateral 1-1 === Each of the multiple elements including the learning unit, generation unit, and advertisement production unit described above is realized, for example, by at least one of the smart device 14 and the data processing device 12. For example, the learning unit is realized by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the generation unit is realized by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the advertisement production unit is realized by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-2 === Each of the multiple elements including the learning unit, generation unit, and advertisement production unit described above is realized, for example, by at least one of the smart glasses 214 and the data processing device 12. For example, the learning unit is realized by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the generation unit is realized by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the advertisement production unit is realized by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-3 === Each of the multiple elements including the learning unit, generation unit, and advertisement production unit described above is realized, for example, by at least one of the headset type terminal 314 and the data processing device 12. For example, the learning unit is realized by the control unit 46A of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the generation unit is realized by the control unit 46A of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the advertisement production unit is realized by the control unit 46A of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-4 === Each of the multiple elements including the learning unit, generation unit, and advertisement production unit described above is realized, for example, by at least one of the robot 414 and the data processing device 12. For example, the learning unit is realized by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the generation unit is realized by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the advertisement production unit is realized by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12.
[0092] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0093] When generating a voice clone, the generator can estimate the user's emotion and adjust the tone and tempo of the voice clone based on the estimated emotion. For example, if the user is relaxed, the generator can generate a voice clone with a relaxed tone and a slow tempo. If the user is excited, the generator can generate a voice clone with an energetic tone and a fast tempo. Furthermore, if the user is sad, the generator can generate a voice clone with a calm tone and a slow tempo. This allows for more personalized advertising by generating a voice clone according to the user's emotion.
[0094] When learning voice data, the learning unit can also take into account the health status of the voice actor or character. For example, if the voice actor has a cold, that voice data can be excluded during learning. Also, if the voice actor is tired, that voice data can be filtered during learning. Furthermore, it can prioritize learning of voice data recorded by the voice actor in optimal condition. This allows for learning that takes into account the health status of the voice, making it possible to generate higher quality voice clones.
[0095] When generating a voice clone, the generation unit can estimate the user's emotion and adjust the emotional expression of the voice clone based on the estimated emotion. For example, if the user is happy, the generation unit can generate a voice clone that expresses the emotion of joy. If the user is angry, the generation unit can also generate a voice clone that expresses the emotion of anger. Furthermore, if the user is sad, the generation unit can also generate a voice clone that expresses the emotion of sadness. This allows for the creation of more emotionally rich advertisements by generating voice clones with emotional expressions that correspond to the user's emotions.
[0096] The advertising production department can also use the generated voice clones to customize the advertisement according to the target audience. For example, for an advertisement targeted at younger generations, the generated voice clone can be used to provide narration in a casual tone. For an advertisement targeted at older people, the generated voice clone can be used to provide narration in a more subdued tone. Furthermore, for advertisements targeted at specific regions, voice clones incorporating the dialect or accent of that region can be used. This allows for customization according to the target audience, maximizing the effectiveness of the advertisement.
[0097] When generating a voice clone, the generation unit can estimate the user's emotion and adjust the intonation of the voice clone based on the estimated emotion. For example, if the user is relaxed, the generation unit can generate a voice clone with a gentle intonation. If the user is excited, the generation unit can generate a voice clone with a strong intonation. Furthermore, if the user is sad, the generation unit can generate a voice clone with a calm intonation. This allows for the creation of more natural advertisements by generating a voice clone with an intonation that matches the user's emotion.
[0098] The advertising production department can use the generated voice clones to add sound effects according to the ad scenario. For example, in action scenes, echo and reverb can be added to the generated voice clones to enhance realism. In emotional scenes, soft effects can be added to the generated voice clones to enhance emotions. Furthermore, in comedy scenes, comical effects can be added to the generated voice clones to elicit laughter. In this way, adding sound effects according to the scenario can enhance the appeal of the advertisement.
[0099] The learning unit can also estimate the user's emotions when learning voice data and select learning data based on the estimated emotions. For example, if the user is relaxed, the learning unit can prioritize learning relaxed voice data. Also, if the user is excited, the learning unit can prioritize learning energetic voice data. Furthermore, if the user is sad, the learning unit can prioritize learning calm voice data. This allows for more effective voice clone generation by selecting learning data according to the user's emotions.
[0100] The generation unit can also adjust the texture of the voice of the voice actor or character when generating a voice clone. For example, if the voice actor has a soft voice, the sound quality of the voice clone can be adjusted to reproduce that texture. Also, if the voice actor has a husky voice, the sound quality of the voice clone can be adjusted to reproduce that texture. Furthermore, if the voice actor has a clear voice, the sound quality of the voice clone can be adjusted to reproduce that texture. In this way, by adjusting the texture of the voice, a more realistic voice clone can be generated.
[0101] The advertising production department can use the generated voice clones to add music and sound effects according to the theme of the advertisement. For example, for an emotional advertisement, emotional music can be added to the generated voice clones. For an action advertisement, powerful sound effects can be added to the generated voice clones. Furthermore, for a comedy advertisement, comical music and sound effects can be added to the generated voice clones. In this way, the appeal of the advertisement can be enhanced by adding music and sound effects according to the theme.
[0102] When generating a voice clone, the generation unit can estimate the user's emotion and adjust the volume of the voice clone based on the estimated emotion. For example, if the user is relaxed, the generation unit can generate a voice clone with a gentle volume. If the user is excited, the generation unit can generate a voice clone with a powerful volume. Furthermore, if the user is sad, the generation unit can generate a voice clone with a calm volume. This allows for the creation of more natural advertisements by generating a voice clone with a volume that corresponds to the user's emotion.
[0103] The processing flow of the second embodiment will be briefly explained below.
[0104] Step 1: The learning unit studies the voice data of voice actors and characters. For example, the learning unit analyzes in detail the tone, rhythm, and pronunciation characteristics of the voice to understand their characteristics. Step 2: The generator generates a voice clone based on the data learned by the trainer. For example, the generator uses a generation AI to generate a high-quality voice clone. The generated voice clone has a quality that is almost indistinguishable from the original voice. Step 3: The advertising production department uses the voice clone generated by the generation department to create an advertisement. For example, the advertising production department uses the generated voice clone to create the narration and character lines of the advertisement. This makes it possible to learn voice data, generate voice clones, and streamline the process of creating advertisements.
[0105] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0106] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0107] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0108] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0109] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0110] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0111] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0112] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0113] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0114] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0115] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0116] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0117] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0118] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0119] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0120] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0121] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0122] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0123] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0124] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0125] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0126] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0127] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0128] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0129] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0130] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0131] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0132] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0133] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0134] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0135] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the identification processing unit 290 using these models.
[0136] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0137] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0138] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0139] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0140] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0141] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0142] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0143] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0144] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0145] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0146] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0147] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0148] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0149] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0150] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0151] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0152] In the robot 414, the processor 46 performs the identification process. The storage 50 stores the identification program 60. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as the control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 also has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform the same process as the identification processing unit 290 using these models.
[0153] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0154] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0155] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0156] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0157] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0158] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0159] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[0160] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[0161] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[0162] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[0163] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[0164] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[0165] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.
[0166] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[0167] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0168] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[0169] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[0170] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[0171] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[0172] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[0173] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.
[0174] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0175] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0176] [Explanation of symbols]
[0177] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot
Claims
1. a learning unit that learns voice data; a generation unit that generates a voice clone based on the data learned by the learning unit; an advertisement production unit that produces advertisements using the voice clones generated by the generation unit; Equipped with A system characterized by:
2. The learning unit Analyzing voice data of voice actors and characters in specific ways to understand their characteristics The system of claim 1 .
3. The generation unit Generate highly accurate voice clones based on trained data The system of claim 1 .
4. The advertising production department Use the generated voice clones to create narration or character dialogue for advertisements The system of claim 1 .
5. The generation unit Generate voice clones in a short time The system of claim 1 .
6. The advertising production department Use generated voice clones to shorten ad production timelines The system of claim 1 .
7. The learning unit The system estimates the user's emotions and adjusts the timing of learning the voice data based on the estimated user emotions. The system of claim 1 .
8. The learning unit When learning voice data, we analyze the past performance history of voice actors and characters to select the appropriate learning method. The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A