System

The system addresses the challenges of generating and sharing custom voice messages by using natural language processing and machine learning to create and share high-quality voice data in the voice of specific celebrities or characters, ensuring appropriateness and user convenience.

JP2026017353APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118135
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Conventional methods struggle to generate custom voice messages that imitate specific celebrities or characters, lack efficient means for reviewing phrase appropriateness, and provide high-quality voice data in real time.

Method used

A system that includes input, transmission, review, voice generation, and sharing means, utilizing natural language processing and machine learning algorithms to create and share customized voice messages, ensuring appropriateness and quality.

Benefits of technology

Enables efficient creation and sharing of high-quality custom voice messages in the voice of specific celebrities or characters, ensuring appropriateness and user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017353000001_ABST
    Figure 2026017353000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system is provided with an input means for inputting a phrase by a user, a transmission means for transmitting the inputted phrase to a server, an examination means for automatically examining the transmitted phrase, a voice generation means for converting the phrase passing the examination into a designated voice, and a transmission means for transmitting the generated voice data to the terminal of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Using conventional methods, it has been difficult to generate custom voice messages that imitate the voice of a specific celebrity or character. Furthermore, there is a lack of efficient means for reviewing whether the generated phrases are appropriate, which means that the phrases entered by the user may be inappropriate. Furthermore, technology for providing high-quality voice data in real time is still immature, so a method is needed for users to quickly and reliably receive customized messages. [Means for solving the problem]

[0005] The present invention provides a system that reproduces phrases entered by a user in the voice of a specific celebrity or character. Specifically, the system includes an input means for a user to input a phrase, a transmission means for transmitting the entered phrase to a server, a review means for automatically reviewing the submitted phrases, a voice generation means for converting phrases that pass the review into a specified voice, and a transmission means for transmitting the generated voice data to a user's terminal. Furthermore, the review means evaluates the appropriateness of the phrase using natural language processing, and the voice generation means uses a machine learning algorithm to imitate the voice of a specific target. The system also includes a means for storing the generated voice data in a database and providing it in a user-accessible format, and a sharing means for enabling a user to share the generated voice message on social media or a messaging application.

[0006] "User" means any individual or entity that generates and uses custom voice messages.

[0007] "Input means" refers to a device or interface through which a user inputs a phrase.

[0008] "Transmission means" refers to the technology or protocol used to transmit input data to the server.

[0009] "Server" means a computer system on a network that receives, stores, and processes input data.

[0010] "Screening tools" refers to algorithms or programs used to determine whether submitted phrases are appropriate.

[0011] "Natural language processing" refers to the technology that enables computers to understand and process human language.

[0012] "Voice generation means" refers to the technology or process that generates a voice that imitates the voice of a particular talent or character.

[0013] A "machine learning algorithm" refers to a method of using large amounts of data to teach a computer to learn specific patterns and characteristics.

[0014] "Audio Data" means audio information stored in digital form.

[0015] A "database" refers to a system for efficiently storing and managing data.

[0016] "Sharing means" refers to the technology or platform for sharing the generated voice message with other users.

[0017] "Social media" refers to online platforms that allow users to exchange information and communicate with each other.

[0018] "Messaging application" refers to software that allows users to send and receive messages in real time. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. An embodiment of this system will be described.

[0041] overview

[0042] The system is primarily comprised of a user, a device, and a server. Users input and send the desired phrase using their device. The sent phrase is received by the server, and after its appropriateness is confirmed through an automatic review, it is generated as voice data in the voice of a designated talent or character using generative AI technology. The generated voice data is sent to the user's device, where the user can play or share it.

[0043] System Details

[0044] User Interface

[0045] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily input the desired phrase and send a request to generate a voice message.

[0046] Data transmission and reception

[0047] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character, typically as an HTTP POST request.

[0048] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0049] Automated Review Process

[0050] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0051] Speech Production Process

[0052] Phrases that pass the screening are input by the server into a voice generation means, where a generative AI model is used to generate voice data that mimics the voice of the talent or character selected by the user. A machine learning algorithm has been pre-trained for this generation.

[0053] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0054] Sending and storing audio data

[0055] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice message at any time later.

[0056] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0057] Share function

[0058] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0059] This allows users to easily and quickly enjoy customized messages in the voice of a specific celebrity or character. As described above, the system of the present invention provides users with high-quality personalized voice messages through a multi-stage process.

[0060] The processing flow will be explained below.

[0061] Step 1:

[0062] Users can open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, and a send button.

[0063] Step 2:

[0064] The user enters the desired phrase in the input field, for example, "Happy birthday, I'll keep on supporting you!"

[0065] Step 3:

[0066] The user selects a voice for a talent or character and clicks the send button, which sends the entered phrase and selection from the device to the server.

[0067] Step 4:

[0068] The device sends the user's input data to the server via an HTTP POST request, which includes the user's phrase and the selected talent or character information.

[0069] Step 5:

[0070] The server processes the received request and stores it in a database, which contains the user's phrases and selections.

[0071] Step 6:

[0072] The server puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrases.

[0073] Step 7:

[0074] If the automated review process detects an inappropriate phrase, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[0075] Step 8:

[0076] The server then feeds the selected phrases into a generative AI model, which uses machine learning algorithms to mimic the voice of the designated talent or character.

[0077] Step 9:

[0078] The generative AI model converts phrases into speech data in a specified voice. For example, the phrase "Happy birthday, I'll always support you!" is generated as speech data in the selected voice.

[0079] Step 10:

[0080] The generated voice data is then sent back to the server and stored in a database, allowing the user to access the voice data later.

[0081] Step 11:

[0082] The server generates a link or file URL to send the audio data to the user's device.

[0083] Step 12:

[0084] The server sends the generated link or file URL to the device, which receives it and displays it in its user interface.

[0085] Step 13:

[0086] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[0087] Step 14:

[0088] Users can share the generated voice message on social media or messaging applications, for example, using LINE or Twitter.

[0089] Example 1

[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0091] With conventional technology, it was difficult for users to properly review and generate high-quality voice messages using the voices of specific celebrities or characters. Furthermore, there was a lack of a mechanism for evaluating the appropriateness of phrases or a way to efficiently share the generated voice data, which hindered user convenience.

[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0093] In this invention, the server includes an input means for a user to input phrases, a transmission means for transmitting the input phrases to the server, a review means for automatically reviewing the transmitted phrases, a voice generation means for converting phrases that pass the review into specified voice, a storage means for saving the generated voice data in a database and generating a link or file URL for the saved voice data, a transmission means for sending the generated link or file URL to the user's terminal, and a playback means for the user to play or download the generated voice data. This makes it possible to efficiently create and share high-quality custom voice messages while ensuring the appropriateness of the phrases.

[0094] "User" means an individual who uses the System to generate custom voice messages using the voice of a particular talent or character.

[0095] "Input means" refers to the interface used by the user to input phrases, including a browser or dedicated application.

[0096] "Submission method" refers to the mechanism by which the entered phrase is sent to the server, typically as an HTTP POST request.

[0097] "Reviewing means" refers to a function for evaluating the appropriateness of submitted phrases. Specifically, it uses natural language processing (NLP) technology to automatically determine the appropriateness of phrases and whether they contain prohibited words.

[0098] "Voice generation means" refers to the technology used to convert the phrases that pass the screening process into a specified voice, using machine learning algorithms to imitate the voice of a specific talent or character.

[0099] "Storage Means" refers to the mechanism for storing the generated audio data in a database and generating a link or file URL to it.

[0100] "Playback means" refers to the function of sending the generated link or file URL to the user's device so that the user can play or download it.

[0101] "Database" means a system for storing transmitted phrases and generated speech data so that they can be accessed at a later time.

[0102] "Link" refers to the URL through which a user can access the generated audio data. By clicking on this link, the user can access the audio data.

[0103] "Machine learning algorithms" refer to programs trained to mimic the voices of specific talent or characters, also known as generative AI models.

[0104] "Natural Language Processing (NLP)" refers to techniques for assessing the appropriateness of phrases. It is used to detect profanity and evaluate the appropriateness of expressions.

[0105] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. The following describes an embodiment of this system.

[0106] System Configuration

[0107] This system mainly consists of users, terminals, and servers.

[0108] Hardware and software used

[0109] A terminal is a computing device (such as a personal computer, smartphone, or tablet) for user input, on which a browser or a dedicated application is installed.

[0110] The server has the computing resources to perform all functions such as receiving data, processing, generating audio, storing data in a database, generating links, etc. The server includes a web server, a database server, and an AI server on which the generative AI model runs.

[0111] A generative AI model is software that uses machine learning algorithms to mimic the voice of a specified talent or character.

[0112] Use software that uses natural language processing (NLP) techniques to determine the appropriateness of a phrase.

[0113] Processing flow and specific examples

[0114] Users can enter the phrase "Happy birthday, I'll keep cheering for you!" and select a specific talent or character. This can be done using the device's browser or a dedicated application.

[0115] The device sends the entered phrase and the selected talent or character information to the server as an HTTP POST request, typically using the HTTPS protocol.

[0116] The server processes the received request and stores it in a database, including the phrase entered by the user and the selected talent or character.

[0117] The server then automatically reviews the stored phrases using natural language processing (NLP) techniques, which evaluates whether the phrases are appropriate and whether they contain any banned words.

[0118] Phrases that pass the screening are input by the server into a voice generation means, which uses a generative AI model to generate voice data in the voice of the talent or character selected by the user. The generative AI model is pre-trained with a machine learning algorithm.

[0119] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0120] The generated audio data is returned to the server and stored in a database. The server then generates a link or file URL for the saved audio data and sends it to the user's device.

[0121] The device displays the link or URL sent from the server in the user interface, and the user can click the link to play or download the audio data.

[0122] Finally, users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0123] This allows users to efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0125] Step 1:

[0126] The user accesses the interface by opening a browser or a dedicated application on their device. The user interface displays a phrase input field, a talent or character selection menu, and a send button. The user inputs the phrase "Happy birthday, I'll keep on cheering for you!" and selects a specific talent or character. The input is captured as keystrokes on the device.

[0127] Step 2:

[0128] The device creates an HTTP POST request based on the phrase entered by the user and the selected talent or character information. This request includes the text data entered by the user and the selected information. The created request is sent to the server using the HTTPS protocol. The input is the user's selected information, and the output is the sent HTTP request.

[0129] Step 3:

[0130] The server receives the incoming HTTP POST request and processes its contents. It extracts the phrase and talent or character information contained in the request and stores it in a database. Specifically, it executes an insert query to the database to create the corresponding record. The input is the HTTP request data, and the output is the updated database result.

[0131] Step 4:

[0132] The server then puts the stored phrases through an automated review process. The review process uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate. Here, the phrase is input into an NLP model to determine its appropriateness and the presence of prohibited words. For example, a phrase such as "happy birthday" would be reviewed. The input is the phrase read from the database, and the output is the review result.

[0133] Step 5:

[0134] Phrases that pass the screening are input by the server into a voice generation means. A generative AI model is used to generate voice data in the voice of a talent or character selected by the user. A machine learning algorithm can then be used to convert this phrase into the specified voice. The input is the phrase that passed the screening and the selected voice model, and the output is the generated voice data.

[0135] Step 6:

[0136] The generated audio data is returned to the server and stored in a database. The stored audio data includes the newly generated audio file and its link information. The input is the generated audio data, and the output is the updated database.

[0137] Step 7:

[0138] The server generates a link or file URL for the audio data stored in the database and sends it to the user's device. Specifically, it generates a link to the stored audio data as an HTTP response and sends it to the device. The input is the audio data record in the database, and the output is the HTTP response.

[0139] Step 8:

[0140] The terminal displays the links and URLs sent from the server in the user interface. The user can click the presented links to play or download the generated audio data. The input is the response data from the server, and the output is the updated result of the user interface.

[0141] Step 9:

[0142] Users share the generated audio message directly from their device to social media or messaging applications, for example by posting the generated audio link on social media platforms such as LINE or Twitter. The input is the generated link, and the output is the social media post.

[0143] In this way, users can efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0144] (Application example 1)

[0145] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0146] Improving the user experience is a key challenge for modern food delivery services. In particular, there is a need for a fun way for users to track the status of their orders. Traditional notification methods are mainly text-based, which can be monotonous for users. Therefore, there is a need for a way to significantly improve the user experience by generating custom voice messages in the voices of specific celebrities or characters to notify delivery status.

[0147] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0148] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, message generation means for generating a specific status notification message using the generated voice data, and playback means for enabling the user to play or share the specific status notification message. This allows the user to receive delivery status notifications in the voice of a specific celebrity or character, and to enjoy checking the progress of their delivery.

[0149] A "user" is a person who can utilize the system to enter phrases to generate and play custom voice messages.

[0150] A "phrase" is text that a user enters to be generated as a voice message in the voice of a particular talent or character.

[0151] "Input means" refers to the interface and device that a user uses to input a phrase.

[0152] "Transmission means" is a communication means having the function of transmitting the input phrase to the server.

[0153] "Screening Measures" are the processes and functions for automatically screening submitted phrases and assessing their appropriateness.

[0154] "Voice generation means" refers to the technology and algorithms used to convert phrases that pass screening into the voice of a specific talent or character.

[0155] "Audio data" refers to a digital audio file generated by an audio generating means.

[0156] The "message generator" is a system for constructing a specific status notification message using the generated voice data.

[0157] "Playback means" refers to the device and interface that allows a user to play or share a generated voice message.

[0158] This invention is a system that allows users to receive specific status notification messages in the voice of a specific celebrity or character. This system is mainly composed of a user, a terminal, a server, and a database.

[0159] User Interface

[0160] Users access the system through a smartphone application, which includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily enter the desired phrase and send a request to the server to generate a voice message.

[0161] Data transmission and reception

[0162] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character. The request is typically sent as an HTTP POST request. The server receives this request and first stores it in a database. The stored data contains the information needed for subsequent processing.

[0163] Automated Review Process

[0164] The server includes a reviewing means for automatically reviewing received phrases. The reviewing means uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate and whether it contains prohibited words. If the phrase contains inappropriate language, an error message is generated.

[0165] Speech Production Process

[0166] Phrases that pass the screening are input into the voice generation means by the server. A generative AI model is used to generate voice data in the voice of the talent or character selected by the user. A machine learning algorithm is pre-trained for generation. For example, if a user sends the phrase "Thank you for your order! Your food is being prepared," the phrase is generated in the voice of a specific talent.

[0167] Message Generation Process

[0168] The generated voice data is configured as a specific situation notification message, such as "Your food is ready! We'll deliver it to you shortly."

[0169] Sending and storing audio data

[0170] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice message at any time later. The server generates a link or file URL to send the voice data to the user's device. This link or URL is sent to the device and displayed in the user interface. The user can click it to play or download the voice data.

[0171] Share function

[0172] Users can share the generated voice messages directly from their devices to social media or messaging applications, such as LINE or Twitter, with friends and family. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters.

[0173] Prompt Sentence Examples

[0174] For example, use the following prompt:

[0175] "Thank you for your order! Your food is ready."

[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0177] Step 1:

[0178] The user inputs a phrase through the application. They enter the phrase in the input field, select a talent or character, and press the send button to complete the input. The notification message entered here is the user's desired message.

[0179] Step 2:

[0180] The device sends a request to the server containing the entered phrase and information about the selected talent or character. This request is sent as an HTTP POST request, and the server receives it. The received data includes the phrase and character information.

[0181] Step 3:

[0182] The server stores the received phrase in a database. The stored data includes information such as the user ID, phrase, selected characters, and timestamp. This preserves the information necessary for subsequent processing.

[0183] Step 4:

[0184] The server then puts the stored phrases through an automated review process. It uses natural language processing (NLP) techniques to determine whether the phrase is appropriate and whether it contains any forbidden words. If it does contain any inappropriate language, it generates an error message and sends it back to the user. This step takes the phrase as input and the review result as output.

[0185] Step 5:

[0186] Phrases that pass the screening are input into a voice generation means. The server uses a machine learning algorithm to convert the phrase into voice data in the voice of the designated talent or character. A generative AI model is responsible for this process. Here, the input is the passed phrase and character information, and the output is voice data.

[0187] Step 6:

[0188] The generated audio data is then sent back to the server and stored in a database. The stored data includes the URL of the audio file and metadata. The user receives a link or file URL as the information needed to access this data.

[0189] Step 7:

[0190] The server uses the generated voice data to generate a specific status notification message and sends it to the user's device. An example of a notification message is "Your food is ready! We'll deliver it to you shortly." The input here is the voice data and status information, and the output is the generated notification message.

[0191] Step 8:

[0192] The user receives a notification message on their device and plays or shares it. By pressing the play button on the application interface, the audio message is played. By pressing the share button, the message can be shared on social media such as LINE or Twitter. Here, the notification message is the input, and audio playback or sharing is the output.

[0193] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0194] The present invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. The following describes in detail the embodiments of the invention.

[0195] overview

[0196] The system consists of a user, a device, a server, and an emotion engine. Users use their device to input and send a desired phrase. The sent phrase is received by the server, and after its appropriateness is confirmed through automatic review, it is generated as voice data using AI technology in a voice tone that corresponds to the user's emotion recognized by the emotion engine. The generated voice data is sent to the user's device, where the user can play or share it.

[0197] System Details

[0198] User Interface

[0199] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone permission screen. This allows users to easily input the desired phrase and send a request to generate a voice message.

[0200] Data transmission and reception

[0201] The device sends a request to the server, typically as an HTTP POST request, containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine.

[0202] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0203] Automated Review Process

[0204] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0205] Emotion Recognition Process

[0206] Phrases that pass the screening process are sent to an emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[0207] Speech Production Process

[0208] The emotion information recognized by the emotion engine is sent to a server and linked to a voice generation means. Phrases and emotion information that pass the screening are input into a generative AI model. This AI model uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[0209] For example, if a user sends the phrase "Happy birthday, I'll keep on supporting you!" and the emotion engine recognizes a positive emotion, it will generate a voice with a bright tone that reflects that emotion.

[0210] Sending and storing audio data

[0211] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[0212] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0213] Share function

[0214] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0215] This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[0216] The processing flow will be explained below.

[0217] Step 1:

[0218] Users open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, a send button, and a camera and microphone permission screen.

[0219] Step 2:

[0220] The user inputs a desired phrase into the input field. For example, the user inputs a phrase such as "Happy birthday, I'll continue to support you!"

[0221] Step 3:

[0222] The user selects the voice of the talent or character, grants permission to access the camera and microphone, and then clicks the send button, which sends the entered phrase, selection information, and image and audio data from the device to the server.

[0223] Step 4:

[0224] The device sends input data and image and audio data required for emotion recognition to the server via an HTTP POST request. This request includes the user's phrase, information about the selected talent or character, and data used by the emotion engine.

[0225] Step 5:

[0226] The server processes the received data and first stores it in a database, which contains the information needed for subsequent processing.

[0227] Step 6:

[0228] The server then puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but if it contains inappropriate language, an error message is generated.

[0229] Step 7:

[0230] If an inappropriate phrase is detected during the automated review process, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[0231] Step 8:

[0232] Phrases that pass the screening are sent by the server to the emotion engine, which estimates the user's emotions from the images and audio data captured by the user's camera and microphone.

[0233] Step 9:

[0234] The emotion engine uses facial recognition and voice analysis technologies to estimate emotions from the user's facial expressions and tone of voice. For example, a smile is recognized as a positive emotion, while tears are recognized as a negative emotion.

[0235] Step 10:

[0236] The recognized emotion information is sent to a server and integrated with the speech generation means, which provides the emotion information to the speech generation means so that phrases are generated with a tone and inflection appropriate to the emotion.

[0237] Step 11:

[0238] The phrases and emotional information that pass the screening process are then fed into a generative AI model, which uses machine learning algorithms to generate audio data that mimics the voice of the designated talent or character, adding tone and inflection that corresponds to the recognized emotion.

[0239] Step 12:

[0240] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[0241] Step 13:

[0242] The server generates a link or file URL to send the audio data to the user's device, which then displays the link or URL in a user interface.

[0243] Step 14:

[0244] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[0245] Step 15:

[0246] Users can share the generated voice message directly from their device to social media or messaging applications, for example, using LINE or Twitter.

[0247] Example 2

[0248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0249] Conventional voice message generation systems have difficulty generating voice messages that reflect the user's emotions, preventing them from providing a personalized experience. Furthermore, the process for evaluating the appropriateness of phrases is insufficient, which can result in inappropriate phrases being generated. Furthermore, when imitating the voices of specific celebrities or characters, natural tones and intonations are often lacking.

[0250] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0251] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for recognizing emotions, and means for adjusting the tone and intonation of the voice in accordance with the emotion recognized by the emotion recognition means. This makes it possible to generate a personalized voice message that reflects the user's emotions and provide a more natural and appropriate voice message.

[0252] An "input means" is a device or software that provides an interface for a user to input a phrase.

[0253] "Transmission means" refers to a device or software that has the function of transmitting phrases and information entered by the user to the server.

[0254] "Screening means" refers to a device or software that has the function of automatically screening transmitted phrases and evaluating their appropriateness and the presence of prohibited words.

[0255] "Speech generation means" refers to a device or software that converts phrases that have passed the screening into a specified voice and generates voice data.

[0256] "Emotion recognition means" refers to devices or software that use algorithms to recognize emotions from the user's facial expressions and voice via the user's camera or microphone.

[0257] The "tone and intonation adjusting means" refers to a device or software that has the function of adjusting the tone and intonation of the generated voice in accordance with the user's emotion recognized by the emotion recognition means.

[0258] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. Components of the system include a user, a terminal, a server, and the emotion engine.

[0259] A user inputs a desired phrase using their own device and sends it. A browser or a dedicated application is installed on the device, which functions as a user interface. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone access permission screen.

[0260] The device generates a request containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine, and sends this request to the server as an HTTP POST request.

[0261] The server receives this request and stores it in a database. The stored data contains information necessary for subsequent processing. The server also subjects the received phrases to an automated review process. This review process uses natural language processing (NLP) techniques to evaluate the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but an error message is generated if it contains inappropriate language.

[0262] Phrases that pass the screening process are sent to the emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[0263] The emotion information recognized by the emotion engine is sent to a server, which then connects to a voice generator. The server then inputs the selected phrases and emotion information into a generative AI model, which uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[0264] As an example of a specific prompt, let's say the user inputs the phrase "Happy birthday! Stay healthy!" in Talent A's voice while smiling. The emotion engine recognizes the user's smile as a positive emotion, and generates a bright-toned voice that reflects that emotion. Here's a concrete example of how this prompt can be input into the generative AI model:

[0265] "Generate the phrase 'Happy birthday! Stay healthy!' in Talent A's voice based on the positive emotions expressed by smiling users."

[0266] The generated audio data is returned to the server and stored in a database. This storage allows the user to access the audio data at any time later. The server generates a link or file URL to send the audio data to the user's device, and this link or URL is sent to the device. The user can play or download the audio data by clicking the link or URL displayed in the user interface.

[0267] Users can also share the generated voice messages directly from their devices on social media and messaging applications. For example, they can share messages with friends and family using LINE or Twitter. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[0268] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0269] Step 1: User enters phrase

[0270] The user accesses the system via the device's browser or a dedicated application. The user selects a phrase to enter in the phrase input field displayed in the user interface and a talent or character from a selection menu. Specifically, the user enters the phrase "Happy birthday! Stay healthy!" and selects Talent A. The entered data is stored internally on the device as a generated prompt.

[0271] input:

[0272] Phrase: "Happy birthday! Stay healthy!"

[0273] Talent or Character: Talent A

[0274] output:

[0275] Generated prompt data

[0276] Step 2: Generate and send a request

[0277] The device generates request data using the entered phrase and information about the selected talent or character, and sends it to the server as an HTTP POST request.

[0278] input:

[0279] Prompt Data

[0280] Data processing and calculation:

[0281] Generating an HTTP POST Request

[0282] output:

[0283] Sending requests to the server

[0284] Step 3: Data received and stored by the server

[0285] The server receives the request sent from the device and stores it in a database, including phrases and talent or character information.

[0286] input:

[0287] Send Request

[0288] Data processing and calculation:

[0289] Saving to a database

[0290] output:

[0291] Phrases and talent information stored in the database

[0292] Step 4: Automated Review Process

[0293] The server automatically reviews stored phrases for appropriateness using natural language processing (NLP) techniques: for example, "Happy Birthday" is deemed appropriate, while inappropriate phrases generate an error message.

[0294] input:

[0295] Phrases in the database

[0296] Data processing and calculation:

[0297] Appropriateness assessment using natural language processing

[0298] output:

[0299] Review result (appropriate or inappropriate)

[0300] Step 5: Emotion Recognition Process

[0301] Phrases deemed appropriate are sent to the emotion engine, which recognizes emotions from facial expressions and tone of voice via the user's camera or microphone. For example, if the user is smiling, it is recognized as a positive emotion.

[0302] input:

[0303] Phrase (reviewed)

[0304] User facial expression data

[0305] User's voice tone

[0306] Data processing and calculation:

[0307] Applying emotion recognition algorithms

[0308] output:

[0309] Recognized emotional information

[0310] Step 6: The speech generation process

[0311] Based on the emotional information from the emotion engine and the screened phrases, the server uses a generative AI model to generate audio in the voice of the designated talent or character, with tone and inflection adjusted according to the recognized emotion.

[0312] input:

[0313] phrase

[0314] Recognized emotional information

[0315] Data processing and calculation:

[0316] Speech generation using generative AI models

[0317] output:

[0318] Generated audio data

[0319] Step 7: Send and save the audio data

[0320] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice data at any time later. The server generates a link or file URL to send the voice data to the user's device and sends this link to the user.

[0321] input:

[0322] Generated audio data

[0323] Data processing and calculation:

[0324] Saving to a database

[0325] Generate a download link or URL

[0326] output:

[0327] Links or URLs sent to users

[0328] Step 8: Sharing Function

[0329] Users receive the generated audio message and can click a link or URL to play or download the audio data, or share the audio message with friends and family via social media or messaging applications.

[0330] input:

[0331] Links or URLs sent to users

[0332] Data processing and calculation:

[0333] Download or play audio data

[0334] Share audio messages

[0335] output:

[0336] Playing audio data

[0337] Share audio messages

[0338] (Application example 2)

[0339] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0340] In modern advertising, sending unilateral messages without considering user emotions cannot be expected to be effective. Furthermore, uniform voice messages have difficulty attracting user attention, and a lack of personalization leads to reduced engagement with advertising. Therefore, there is a need for a system that can grasp a user's emotional state in real time and generate and deliver personalized advertising messages with appropriate voice tones.

[0341] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for the user to input a phrase, a transmission means for transmitting the input phrase to the server, an examination means for automatically examining the transmitted phrase, an emotion recognition means for recognizing an emotion, a voice generation means for converting the phrase into a specified voice with an audio tone corresponding to the emotion recognized by the emotion recognition means, and a transmission means for transmitting the generated voice data to the user's terminal and playing it back. This makes it possible to generate and deliver a personalized voice message that takes the user's emotions into consideration.

[0342] "User" means a person or entity that utilizes the system to input phrases and generate voice messages.

[0343] "Input means" refers to a device or software function that allows a user to input a desired phrase.

[0344] "Transmission means" refers to a device or software function for transmitting the input phrase to the server.

[0345] A "screening tool" is a process or algorithm that automatically determines the appropriateness of a submitted phrase.

[0346] "Emotion recognition means" refers to a device or software function that reads and analyzes a user's emotions through a camera or microphone.

[0347] A "voice generation means" is a device or software function that converts phrases into voice data based on a specified voice or emotional tone.

[0348] The "server" is a central processing unit that processes various data and transmits audio data and control information to client terminals.

[0349] A "terminal" is a device (such as a smartphone or PC) that a user directly operates and interacts with the system.

[0350] A "phrase" is a series of words or sentences that a user wishes to convert into a voice message.

[0351] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. This system incorporates an emotion engine that recognizes the user's emotions, enabling the creation of more personalized voice messages.

[0352] composition

[0353] This system consists of a user terminal, a server, an emotion recognition engine, and a voice generation engine. The user terminal is assumed to be a smartphone, but other devices can also be used. The server acts as a central processing unit and manages the emotion recognition and voice generation processes.

[0354] Hardware / Software used

[0355] Hardware:

[0356] Smartphone (iOS, Android)

[0357] software:

[0358] Microsoft Azure's Face API or Google Cloud's Vision AI (emotion recognition)

[0359] OpenAI's GPT-3 (Speech Generation)

[0360] HTTP server (data transmission and reception)

[0361] Operation method and configuration

[0362] 1. User Interface:

[0363] The user launches the app and is asked for camera and microphone permissions, which initiates the emotion recognition process. The user then enters the desired message in the phrase input field and presses send.

[0364] 2. Emotion recognition:

[0365] Data is collected in real time from the camera and microphone on the user's device and sent to the emotion recognition API, where the user's facial expressions and tone of voice are analyzed to obtain emotional information.

[0366] 3. Data transmission:

[0367] The acquired emotion information and the input phrase are sent to the server using an HTTP POST request.

[0368] 4. Automated Review:

[0369] The server receives the submitted phrase and evaluates its appropriateness using natural language processing (NLP) techniques. If it contains inappropriate language, it returns an error message; if it is an appropriate phrase, it proceeds to the next step.

[0370] 5. Speech generation:

[0371] Based on the phrases and emotional information that pass the screening process, a speech generation engine (OpenAI's GPT-3) is used to generate a custom voice message in the voice of the designated talent or character, with the tone and intonation adjusted according to the recognized emotion.

[0372] 6. Data transmission and playback:

[0373] The generated audio data is then sent back to the user's device via the server, where it can be played within the application, and a sharing link is also generated, allowing the user to share the audio message on social media or via messaging apps.

[0374] Specific examples

[0375] For example, if a user is in a good mood and types the phrase "Thank you for your hard work today!", the emotion recognition engine will recognize the positive emotion and generate a message in a bright tone such as "Thank you for your hard work today!"

[0376] Example prompt sentence:

[0377] "Please generate an advertising message in the voice of celebrity X to be played when the user has a happy expression. The message should introduce product Z in a cheerful tone."

[0378] In this way, it becomes possible to generate and deliver personalized voice messages that take into account the user's emotions.

[0379] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0380] Step 1:

[0381] The user launches the application and grants permission to use the camera and microphone. Once this is complete, the system is ready to collect facial and voice data in real time.

[0382] Input: App launch, camera and microphone permissions

[0383] Output: Real-time facial expression and voice data acquisition status

[0384] Step 2:

[0385] The user enters the phrase using an input device and presses the send button, at which point the phrase is saved on the device and ready to be sent in the next step.

[0386] Input: A phrase entered by the user

[0387] Output: The phrase you entered is saved to your device.

[0388] Step 3:

[0389] The user device sends the phrase entered and the facial expression and voice data acquired in real time to the server as an HTTP POST request. The transmitted data includes the phrase entered by the user, facial expression data, and voice data.

[0390] Input: User-entered phrases, facial expressions and voice data captured in real time

[0391] Output: Request data sent to the server

[0392] Step 4:

[0393] The server then puts the received phrases through an automated review process. The review uses natural language processing (NLP) technology to evaluate the appropriateness of the phrase. If the phrase contains inappropriate content, an error message is generated and returned to the terminal. If the phrase is appropriate, the process proceeds to the next step.

[0394] Input: Phrase sent to the server

[0395] Output: Evaluation result for suitability, error message or approval flag

[0396] Step 5:

[0397] The server sends the facial and voice data acquired in real time to the emotion recognition device, which then analyzes the data and recognizes the user's emotions (such as Microsoft Azure's Face API or Google Cloud's Vision AI). The recognized emotion data is then sent to the server.

[0398] Input: Real-time facial and voice data

[0399] Output: Recognized emotion data

[0400] Step 6:

[0401] Based on the phrases and recognized emotion data that have passed the screening process, the server uses a voice generation engine (OpenAI's GPT-3) to generate a custom voice message in the voice of the specified talent or character, taking into account emotional tone during the generation process.

[0402] Input: phrases that passed the screening, recognized emotion data

[0403] Output: Generated audio data

[0404] Step 7:

[0405] The server sends the generated audio data to the user terminal. In this step, the generated audio data and a playback link are sent to the user terminal.

[0406] Input: Generated audio data

[0407] Output: Audio data and playback link sent to the user's device

[0408] Step 8:

[0409] The user's device will then play the received audio data, allowing the user to listen to it. Additionally, the user will have the ability to share the audio message on social media or messaging apps.

[0410] Input: Audio data and playback link sent to the user's device

[0411] Output: Played audio data and a shareable link

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0413] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0415] [Second embodiment]

[0416] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0417] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0418] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0420] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0422] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0423] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0424] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0427] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0428] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. An embodiment of this system will be described.

[0429] overview

[0430] The system is primarily comprised of a user, a device, and a server. Users input and send the desired phrase using their device. The sent phrase is received by the server, and after its appropriateness is confirmed through an automatic review, it is generated as voice data in the voice of a designated talent or character using generative AI technology. The generated voice data is sent to the user's device, where the user can play or share it.

[0431] System Details

[0432] User Interface

[0433] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily input the desired phrase and send a request to generate a voice message.

[0434] Data transmission and reception

[0435] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character, typically as an HTTP POST request.

[0436] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0437] Automated Review Process

[0438] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0439] Speech Production Process

[0440] Phrases that pass the screening are input by the server into a voice generation means, where a generative AI model is used to generate voice data that mimics the voice of the talent or character selected by the user. A machine learning algorithm has been pre-trained for this generation.

[0441] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0442] Sending and storing audio data

[0443] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice message at any time later.

[0444] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0445] Share function

[0446] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0447] This allows users to easily and quickly enjoy customized messages in the voice of a specific celebrity or character. As described above, the system of the present invention provides users with high-quality personalized voice messages through a multi-stage process.

[0448] The processing flow will be explained below.

[0449] Step 1:

[0450] Users can open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, and a send button.

[0451] Step 2:

[0452] The user enters the desired phrase in the input field, for example, "Happy birthday, I'll keep on supporting you!"

[0453] Step 3:

[0454] The user selects a voice for a talent or character and clicks the send button, which sends the entered phrase and selection from the device to the server.

[0455] Step 4:

[0456] The device sends the user's input data to the server via an HTTP POST request, which includes the user's phrase and the selected talent or character information.

[0457] Step 5:

[0458] The server processes the received request and stores it in a database, which contains the user's phrases and selections.

[0459] Step 6:

[0460] The server puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrases.

[0461] Step 7:

[0462] If the automated review process detects an inappropriate phrase, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[0463] Step 8:

[0464] The server then feeds the selected phrases into a generative AI model, which uses machine learning algorithms to mimic the voice of the designated talent or character.

[0465] Step 9:

[0466] The generative AI model converts phrases into speech data in a specified voice. For example, the phrase "Happy birthday, I'll always support you!" is generated as speech data in the selected voice.

[0467] Step 10:

[0468] The generated voice data is then sent back to the server and stored in a database, allowing the user to access the voice data later.

[0469] Step 11:

[0470] The server generates a link or file URL to send the audio data to the user's device.

[0471] Step 12:

[0472] The server sends the generated link or file URL to the device, which receives it and displays it in its user interface.

[0473] Step 13:

[0474] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[0475] Step 14:

[0476] Users can share the generated voice message on social media or messaging applications, for example, using LINE or Twitter.

[0477] Example 1

[0478] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0479] With conventional technology, it was difficult for users to properly review and generate high-quality voice messages using the voices of specific celebrities or characters. Furthermore, there was a lack of a mechanism for evaluating the appropriateness of phrases or a way to efficiently share the generated voice data, which hindered user convenience.

[0480] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0481] In this invention, the server includes an input means for a user to input phrases, a transmission means for transmitting the input phrases to the server, a review means for automatically reviewing the transmitted phrases, a voice generation means for converting phrases that pass the review into specified voice, a storage means for saving the generated voice data in a database and generating a link or file URL for the saved voice data, a transmission means for sending the generated link or file URL to the user's terminal, and a playback means for the user to play or download the generated voice data. This makes it possible to efficiently create and share high-quality custom voice messages while ensuring the appropriateness of the phrases.

[0482] "User" means an individual who uses the System to generate custom voice messages using the voice of a particular talent or character.

[0483] "Input means" refers to the interface used by the user to input phrases, including a browser or dedicated application.

[0484] "Submission method" refers to the mechanism by which the entered phrase is sent to the server, typically as an HTTP POST request.

[0485] "Reviewing means" refers to a function for evaluating the appropriateness of submitted phrases. Specifically, it uses natural language processing (NLP) technology to automatically determine the appropriateness of phrases and whether they contain prohibited words.

[0486] "Voice generation means" refers to the technology used to convert the phrases that pass the screening process into a specified voice, using machine learning algorithms to imitate the voice of a specific talent or character.

[0487] "Storage Means" refers to the mechanism for storing the generated audio data in a database and generating a link or file URL to it.

[0488] "Playback means" refers to the function of sending the generated link or file URL to the user's device so that the user can play or download it.

[0489] "Database" means a system for storing transmitted phrases and generated speech data so that they can be accessed at a later time.

[0490] "Link" refers to the URL through which a user can access the generated audio data. By clicking on this link, the user can access the audio data.

[0491] "Machine learning algorithms" refer to programs trained to mimic the voices of specific talent or characters, also known as generative AI models.

[0492] "Natural Language Processing (NLP)" refers to techniques for assessing the appropriateness of phrases. It is used to detect profanity and evaluate the appropriateness of expressions.

[0493] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. The following describes an embodiment of this system.

[0494] System Configuration

[0495] This system mainly consists of users, terminals, and servers.

[0496] Hardware and software used

[0497] A terminal is a computing device (such as a personal computer, smartphone, or tablet) for user input, on which a browser or a dedicated application is installed.

[0498] The server has the computing resources to perform all functions such as receiving data, processing, generating audio, storing data in a database, generating links, etc. The server includes a web server, a database server, and an AI server on which the generative AI model runs.

[0499] A generative AI model is software that uses machine learning algorithms to mimic the voice of a specified talent or character.

[0500] Use software that uses natural language processing (NLP) techniques to determine the appropriateness of a phrase.

[0501] Processing flow and specific examples

[0502] Users can enter the phrase "Happy birthday, I'll keep cheering for you!" and select a specific talent or character. This can be done using the device's browser or a dedicated application.

[0503] The device sends the entered phrase and the selected talent or character information to the server as an HTTP POST request, typically using the HTTPS protocol.

[0504] The server processes the received request and stores it in a database, including the phrase entered by the user and the selected talent or character.

[0505] The server then automatically reviews the stored phrases using natural language processing (NLP) techniques, which evaluates whether the phrases are appropriate and whether they contain any banned words.

[0506] Phrases that pass the screening are input by the server into a voice generation means, which uses a generative AI model to generate voice data in the voice of the talent or character selected by the user. The generative AI model is pre-trained with a machine learning algorithm.

[0507] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0508] The generated audio data is returned to the server and stored in a database. The server then generates a link or file URL for the saved audio data and sends it to the user's device.

[0509] The device displays the link or URL sent from the server in the user interface, and the user can click the link to play or download the audio data.

[0510] Finally, users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0511] This allows users to efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0512] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0513] Step 1:

[0514] The user accesses the interface by opening a browser or a dedicated application on their device. The user interface displays a phrase input field, a talent or character selection menu, and a send button. The user inputs the phrase "Happy birthday, I'll keep on cheering for you!" and selects a specific talent or character. The input is captured as keystrokes on the device.

[0515] Step 2:

[0516] The device creates an HTTP POST request based on the phrase entered by the user and the selected talent or character information. This request includes the text data entered by the user and the selected information. The created request is sent to the server using the HTTPS protocol. The input is the user's selected information, and the output is the sent HTTP request.

[0517] Step 3:

[0518] The server receives the incoming HTTP POST request and processes its contents. It extracts the phrase and talent or character information contained in the request and stores it in a database. Specifically, it executes an insert query to the database to create the corresponding record. The input is the HTTP request data, and the output is the updated database result.

[0519] Step 4:

[0520] The server then puts the stored phrases through an automated review process. The review process uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate. Here, the phrase is input into an NLP model to determine its appropriateness and the presence of prohibited words. For example, a phrase such as "happy birthday" would be reviewed. The input is the phrase read from the database, and the output is the review result.

[0521] Step 5:

[0522] Phrases that pass the screening are input by the server into a voice generation means. A generative AI model is used to generate voice data in the voice of a talent or character selected by the user. A machine learning algorithm can then be used to convert this phrase into the specified voice. The input is the phrase that passed the screening and the selected voice model, and the output is the generated voice data.

[0523] Step 6:

[0524] The generated audio data is returned to the server and stored in a database. The stored audio data includes the newly generated audio file and its link information. The input is the generated audio data, and the output is the updated database.

[0525] Step 7:

[0526] The server generates a link or file URL for the audio data stored in the database and sends it to the user's device. Specifically, it generates a link to the stored audio data as an HTTP response and sends it to the device. The input is the audio data record in the database, and the output is the HTTP response.

[0527] Step 8:

[0528] The terminal displays the links and URLs sent from the server in the user interface. The user can click the presented links to play or download the generated audio data. The input is the response data from the server, and the output is the updated result of the user interface.

[0529] Step 9:

[0530] Users share the generated audio message directly from their device to social media or messaging applications, for example by posting the generated audio link on social media platforms such as LINE or Twitter. The input is the generated link, and the output is the social media post.

[0531] In this way, users can efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0532] (Application example 1)

[0533] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0534] Improving the user experience is a key challenge for modern food delivery services. In particular, there is a need for a fun way for users to track the status of their orders. Traditional notification methods are mainly text-based, which can be monotonous for users. Therefore, there is a need for a way to significantly improve the user experience by generating custom voice messages in the voices of specific celebrities or characters to notify delivery status.

[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0536] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, message generation means for generating a specific status notification message using the generated voice data, and playback means for enabling the user to play or share the specific status notification message. This allows the user to receive delivery status notifications in the voice of a specific celebrity or character, and to enjoy checking the progress of their delivery.

[0537] A "user" is a person who can utilize the system to enter phrases to generate and play custom voice messages.

[0538] A "phrase" is text that a user enters to be generated as a voice message in the voice of a particular talent or character.

[0539] "Input means" refers to the interface and device that a user uses to input a phrase.

[0540] "Transmission means" is a communication means having the function of transmitting the input phrase to the server.

[0541] "Screening Measures" are the processes and functions for automatically screening submitted phrases and assessing their appropriateness.

[0542] "Voice generation means" refers to the technology and algorithms used to convert phrases that pass screening into the voice of a specific talent or character.

[0543] "Audio data" refers to a digital audio file generated by an audio generating means.

[0544] The "message generator" is a system for constructing a specific status notification message using the generated voice data.

[0545] "Playback means" refers to the device and interface that allows a user to play or share a generated voice message.

[0546] This invention is a system that allows users to receive specific status notification messages in the voice of a specific celebrity or character. This system is mainly composed of a user, a terminal, a server, and a database.

[0547] User Interface

[0548] Users access the system through a smartphone application, which includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily enter the desired phrase and send a request to the server to generate a voice message.

[0549] Data transmission and reception

[0550] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character. The request is typically sent as an HTTP POST request. The server receives this request and first stores it in a database. The stored data contains the information needed for subsequent processing.

[0551] Automated Review Process

[0552] The server includes a reviewing means for automatically reviewing received phrases. The reviewing means uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate and whether it contains prohibited words. If the phrase contains inappropriate language, an error message is generated.

[0553] Speech Production Process

[0554] Phrases that pass the screening are input into the voice generation means by the server. A generative AI model is used to generate voice data in the voice of the talent or character selected by the user. A machine learning algorithm is pre-trained for generation. For example, if a user sends the phrase "Thank you for your order! Your food is being prepared," the phrase is generated in the voice of a specific talent.

[0555] Message Generation Process

[0556] The generated voice data is configured as a specific situation notification message, such as "Your food is ready! We'll deliver it to you shortly."

[0557] Sending and storing audio data

[0558] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice message at any time later. The server generates a link or file URL to send the voice data to the user's device. This link or URL is sent to the device and displayed in the user interface. The user can click it to play or download the voice data.

[0559] Share function

[0560] Users can share the generated voice messages directly from their devices to social media or messaging applications, such as LINE or Twitter, with friends and family. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters.

[0561] Prompt Sentence Examples

[0562] For example, use the following prompt:

[0563] "Thank you for your order! Your food is ready."

[0564] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0565] Step 1:

[0566] The user inputs a phrase through the application. They enter the phrase in the input field, select a talent or character, and press the send button to complete the input. The notification message entered here is the user's desired message.

[0567] Step 2:

[0568] The device sends a request to the server containing the entered phrase and information about the selected talent or character. This request is sent as an HTTP POST request, and the server receives it. The received data includes the phrase and character information.

[0569] Step 3:

[0570] The server stores the received phrase in a database. The stored data includes information such as the user ID, phrase, selected characters, and timestamp. This preserves the information necessary for subsequent processing.

[0571] Step 4:

[0572] The server then puts the stored phrases through an automated review process. It uses natural language processing (NLP) techniques to determine whether the phrase is appropriate and whether it contains any forbidden words. If it does contain any inappropriate language, it generates an error message and sends it back to the user. This step takes the phrase as input and the review result as output.

[0573] Step 5:

[0574] Phrases that pass the screening are input into a voice generation means. The server uses a machine learning algorithm to convert the phrase into voice data in the voice of the designated talent or character. A generative AI model is responsible for this process. Here, the input is the passed phrase and character information, and the output is voice data.

[0575] Step 6:

[0576] The generated audio data is then sent back to the server and stored in a database. The stored data includes the URL of the audio file and metadata. The user receives a link or file URL as the information needed to access this data.

[0577] Step 7:

[0578] The server uses the generated voice data to generate a specific status notification message and sends it to the user's device. An example of a notification message is "Your food is ready! We'll deliver it to you shortly." The input here is the voice data and status information, and the output is the generated notification message.

[0579] Step 8:

[0580] The user receives a notification message on their device and plays or shares it. By pressing the play button on the application interface, the audio message is played. By pressing the share button, the message can be shared on social media such as LINE or Twitter. Here, the notification message is the input, and audio playback or sharing is the output.

[0581] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0582] The present invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. The following describes in detail the embodiments of the invention.

[0583] overview

[0584] The system consists of a user, a device, a server, and an emotion engine. Users use their device to input and send a desired phrase. The sent phrase is received by the server, and after its appropriateness is confirmed through automatic review, it is generated as voice data using AI technology in a voice tone that corresponds to the user's emotion recognized by the emotion engine. The generated voice data is sent to the user's device, where the user can play or share it.

[0585] System Details

[0586] User Interface

[0587] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone permission screen. This allows users to easily input the desired phrase and send a request to generate a voice message.

[0588] Data transmission and reception

[0589] The device sends a request to the server, typically as an HTTP POST request, containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine.

[0590] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0591] Automated Review Process

[0592] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0593] Emotion Recognition Process

[0594] Phrases that pass the screening process are sent to an emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[0595] Speech Production Process

[0596] The emotion information recognized by the emotion engine is sent to a server and linked to a voice generation means. Phrases and emotion information that pass the screening are input into a generative AI model. This AI model uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[0597] For example, if a user sends the phrase "Happy birthday, I'll keep on supporting you!" and the emotion engine recognizes a positive emotion, it will generate a voice with a bright tone that reflects that emotion.

[0598] Sending and storing audio data

[0599] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[0600] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0601] Share function

[0602] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0603] This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[0604] The processing flow will be explained below.

[0605] Step 1:

[0606] Users open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, a send button, and a camera and microphone permission screen.

[0607] Step 2:

[0608] The user inputs a desired phrase into the input field. For example, the user inputs a phrase such as "Happy birthday, I'll continue to support you!"

[0609] Step 3:

[0610] The user selects the voice of the talent or character, grants permission to access the camera and microphone, and then clicks the send button, which sends the entered phrase, selection information, and image and audio data from the device to the server.

[0611] Step 4:

[0612] The device sends input data and image and audio data required for emotion recognition to the server via an HTTP POST request. This request includes the user's phrase, information about the selected talent or character, and data used by the emotion engine.

[0613] Step 5:

[0614] The server processes the received data and first stores it in a database, which contains the information needed for subsequent processing.

[0615] Step 6:

[0616] The server then puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but if it contains inappropriate language, an error message is generated.

[0617] Step 7:

[0618] If an inappropriate phrase is detected during the automated review process, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[0619] Step 8:

[0620] Phrases that pass the screening are sent by the server to the emotion engine, which estimates the user's emotions from the images and audio data captured by the user's camera and microphone.

[0621] Step 9:

[0622] The emotion engine uses facial recognition and voice analysis technologies to estimate emotions from the user's facial expressions and tone of voice. For example, a smile is recognized as a positive emotion, while tears are recognized as a negative emotion.

[0623] Step 10:

[0624] The recognized emotion information is sent to a server and integrated with the speech generation means, which provides the emotion information to the speech generation means so that phrases are generated with a tone and inflection appropriate to the emotion.

[0625] Step 11:

[0626] The phrases and emotional information that pass the screening process are then fed into a generative AI model, which uses machine learning algorithms to generate audio data that mimics the voice of the designated talent or character, adding tone and inflection that corresponds to the recognized emotion.

[0627] Step 12:

[0628] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[0629] Step 13:

[0630] The server generates a link or file URL to send the audio data to the user's device, which then displays the link or URL in a user interface.

[0631] Step 14:

[0632] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[0633] Step 15:

[0634] Users can share the generated voice message directly from their device to social media or messaging applications, for example, using LINE or Twitter.

[0635] Example 2

[0636] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0637] Conventional voice message generation systems have difficulty generating voice messages that reflect the user's emotions, preventing them from providing a personalized experience. Furthermore, the process for evaluating the appropriateness of phrases is insufficient, which can result in inappropriate phrases being generated. Furthermore, when imitating the voices of specific celebrities or characters, natural tones and intonations are often lacking.

[0638] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0639] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for recognizing emotions, and means for adjusting the tone and intonation of the voice in accordance with the emotion recognized by the emotion recognition means. This makes it possible to generate a personalized voice message that reflects the user's emotions and provide a more natural and appropriate voice message.

[0640] An "input means" is a device or software that provides an interface for a user to input a phrase.

[0641] "Transmission means" refers to a device or software that has the function of transmitting phrases and information entered by the user to the server.

[0642] "Screening means" refers to a device or software that has the function of automatically screening transmitted phrases and evaluating their appropriateness and the presence of prohibited words.

[0643] "Speech generation means" refers to a device or software that converts phrases that have passed the screening into a specified voice and generates voice data.

[0644] "Emotion recognition means" refers to devices or software that use algorithms to recognize emotions from the user's facial expressions and voice via the user's camera or microphone.

[0645] The "tone and intonation adjusting means" refers to a device or software that has the function of adjusting the tone and intonation of the generated voice in accordance with the user's emotion recognized by the emotion recognition means.

[0646] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. Components of the system include a user, a terminal, a server, and the emotion engine.

[0647] A user inputs a desired phrase using their own device and sends it. A browser or a dedicated application is installed on the device, which functions as a user interface. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone access permission screen.

[0648] The device generates a request containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine, and sends this request to the server as an HTTP POST request.

[0649] The server receives this request and stores it in a database. The stored data contains information necessary for subsequent processing. The server also subjects the received phrases to an automated review process. This review process uses natural language processing (NLP) techniques to evaluate the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but an error message is generated if it contains inappropriate language.

[0650] Phrases that pass the screening process are sent to the emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[0651] The emotion information recognized by the emotion engine is sent to a server, which then connects to a voice generator. The server then inputs the selected phrases and emotion information into a generative AI model, which uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[0652] As an example of a specific prompt, let's say the user inputs the phrase "Happy birthday! Stay healthy!" in Talent A's voice while smiling. The emotion engine recognizes the user's smile as a positive emotion, and generates a bright-toned voice that reflects that emotion. Here's a concrete example of how this prompt can be input into the generative AI model:

[0653] "Generate the phrase 'Happy birthday! Stay healthy!' in Talent A's voice based on the positive emotions expressed by smiling users."

[0654] The generated audio data is returned to the server and stored in a database. This storage allows the user to access the audio data at any time later. The server generates a link or file URL to send the audio data to the user's device, and this link or URL is sent to the device. The user can play or download the audio data by clicking the link or URL displayed in the user interface.

[0655] Users can also share the generated voice messages directly from their devices on social media and messaging applications. For example, they can share messages with friends and family using LINE or Twitter. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[0656] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0657] Step 1: User enters phrase

[0658] The user accesses the system via the device's browser or a dedicated application. The user selects a phrase to enter in the phrase input field displayed in the user interface and a talent or character from a selection menu. Specifically, the user enters the phrase "Happy birthday! Stay healthy!" and selects Talent A. The entered data is stored internally on the device as a generated prompt.

[0659] input:

[0660] Phrase: "Happy birthday! Stay healthy!"

[0661] Talent or Character: Talent A

[0662] output:

[0663] Generated prompt data

[0664] Step 2: Generate and send a request

[0665] The device generates request data using the entered phrase and information about the selected talent or character, and sends it to the server as an HTTP POST request.

[0666] input:

[0667] Prompt Data

[0668] Data processing and calculation:

[0669] Generating an HTTP POST Request

[0670] output:

[0671] Sending requests to the server

[0672] Step 3: Data received and stored by the server

[0673] The server receives the request sent from the device and stores it in a database, including phrases and talent or character information.

[0674] input:

[0675] Send Request

[0676] Data processing and calculation:

[0677] Saving to a database

[0678] output:

[0679] Phrases and talent information stored in the database

[0680] Step 4: Automated Review Process

[0681] The server automatically reviews stored phrases for appropriateness using natural language processing (NLP) techniques: for example, "Happy Birthday" is deemed appropriate, while inappropriate phrases generate an error message.

[0682] input:

[0683] Phrases in the database

[0684] Data processing and calculation:

[0685] Appropriateness assessment using natural language processing

[0686] output:

[0687] Review result (appropriate or inappropriate)

[0688] Step 5: Emotion Recognition Process

[0689] Phrases deemed appropriate are sent to the emotion engine, which recognizes emotions from facial expressions and tone of voice via the user's camera or microphone. For example, if the user is smiling, it is recognized as a positive emotion.

[0690] input:

[0691] Phrase (reviewed)

[0692] User facial expression data

[0693] User's voice tone

[0694] Data processing and calculation:

[0695] Applying emotion recognition algorithms

[0696] output:

[0697] Recognized emotional information

[0698] Step 6: The speech generation process

[0699] Based on the emotional information from the emotion engine and the screened phrases, the server uses a generative AI model to generate audio in the voice of the designated talent or character, with tone and inflection adjusted according to the recognized emotion.

[0700] input:

[0701] phrase

[0702] Recognized emotional information

[0703] Data processing and calculation:

[0704] Speech generation using generative AI models

[0705] output:

[0706] Generated audio data

[0707] Step 7: Send and save the audio data

[0708] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice data at any time later. The server generates a link or file URL to send the voice data to the user's device and sends this link to the user.

[0709] input:

[0710] Generated audio data

[0711] Data processing and calculation:

[0712] Saving to a database

[0713] Generate a download link or URL

[0714] output:

[0715] Links or URLs sent to users

[0716] Step 8: Sharing Function

[0717] Users receive the generated audio message and can click a link or URL to play or download the audio data, or share the audio message with friends and family via social media or messaging applications.

[0718] input:

[0719] Links or URLs sent to users

[0720] Data processing and calculation:

[0721] Download or play audio data

[0722] Share audio messages

[0723] output:

[0724] Playing audio data

[0725] Share audio messages

[0726] (Application example 2)

[0727] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0728] In modern advertising, sending unilateral messages without considering user emotions cannot be expected to be effective. Furthermore, uniform voice messages have difficulty attracting user attention, and a lack of personalization leads to reduced engagement with advertising. Therefore, there is a need for a system that can grasp a user's emotional state in real time and generate and deliver personalized advertising messages with appropriate voice tones.

[0729] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for the user to input a phrase, a transmission means for transmitting the input phrase to the server, an examination means for automatically examining the transmitted phrase, an emotion recognition means for recognizing an emotion, a voice generation means for converting the phrase into a specified voice with an audio tone corresponding to the emotion recognized by the emotion recognition means, and a transmission means for transmitting the generated voice data to the user's terminal and playing it back. This makes it possible to generate and deliver a personalized voice message that takes the user's emotions into consideration.

[0730] "User" means a person or entity that utilizes the system to input phrases and generate voice messages.

[0731] "Input means" refers to a device or software function that allows a user to input a desired phrase.

[0732] "Transmission means" refers to a device or software function for transmitting the input phrase to the server.

[0733] A "screening tool" is a process or algorithm that automatically determines the appropriateness of a submitted phrase.

[0734] "Emotion recognition means" refers to a device or software function that reads and analyzes a user's emotions through a camera or microphone.

[0735] A "voice generation means" is a device or software function that converts phrases into voice data based on a specified voice or emotional tone.

[0736] The "server" is a central processing unit that processes various data and transmits audio data and control information to client terminals.

[0737] A "terminal" is a device (such as a smartphone or PC) that a user directly operates and interacts with the system.

[0738] A "phrase" is a series of words or sentences that a user wishes to convert into a voice message.

[0739] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. This system incorporates an emotion engine that recognizes the user's emotions, enabling the creation of more personalized voice messages.

[0740] composition

[0741] This system consists of a user terminal, a server, an emotion recognition engine, and a voice generation engine. The user terminal is assumed to be a smartphone, but other devices can also be used. The server acts as a central processing unit and manages the emotion recognition and voice generation processes.

[0742] Hardware / Software used

[0743] Hardware:

[0744] Smartphone (iOS, Android)

[0745] software:

[0746] Microsoft Azure's Face API or Google Cloud's Vision AI (emotion recognition)

[0747] OpenAI's GPT-3 (Speech Generation)

[0748] HTTP server (data transmission and reception)

[0749] Operation method and configuration

[0750] 1. User Interface:

[0751] The user launches the app and is asked for camera and microphone permissions, which initiates the emotion recognition process. The user then enters the desired message in the phrase input field and presses send.

[0752] 2. Emotion recognition:

[0753] Data is collected in real time from the camera and microphone on the user's device and sent to the emotion recognition API, where the user's facial expressions and tone of voice are analyzed to obtain emotional information.

[0754] 3. Data transmission:

[0755] The acquired emotion information and the input phrase are sent to the server using an HTTP POST request.

[0756] 4. Automated Review:

[0757] The server receives the submitted phrase and evaluates its appropriateness using natural language processing (NLP) techniques. If it contains inappropriate language, it returns an error message; if it is an appropriate phrase, it proceeds to the next step.

[0758] 5. Speech generation:

[0759] Based on the phrases and emotional information that pass the screening process, a speech generation engine (OpenAI's GPT-3) is used to generate a custom voice message in the voice of the designated talent or character, with the tone and intonation adjusted according to the recognized emotion.

[0760] 6. Data transmission and playback:

[0761] The generated audio data is then sent back to the user's device via the server, where it can be played within the application, and a sharing link is also generated, allowing the user to share the audio message on social media or via messaging apps.

[0762] Specific examples

[0763] For example, if a user is in a good mood and types the phrase "Thank you for your hard work today!", the emotion recognition engine will recognize the positive emotion and generate a message in a bright tone such as "Thank you for your hard work today!"

[0764] Example prompt sentence:

[0765] "Please generate an advertising message in the voice of celebrity X to be played when the user has a happy expression. The message should introduce product Z in a cheerful tone."

[0766] In this way, it becomes possible to generate and deliver personalized voice messages that take into account the user's emotions.

[0767] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0768] Step 1:

[0769] The user launches the application and grants permission to use the camera and microphone. Once this is complete, the system is ready to collect facial and voice data in real time.

[0770] Input: App launch, camera and microphone permissions

[0771] Output: Real-time facial expression and voice data acquisition status

[0772] Step 2:

[0773] The user enters the phrase using an input device and presses the send button, at which point the phrase is saved on the device and ready to be sent in the next step.

[0774] Input: A phrase entered by the user

[0775] Output: The phrase you entered is saved to your device.

[0776] Step 3:

[0777] The user device sends the phrase entered and the facial expression and voice data acquired in real time to the server as an HTTP POST request. The transmitted data includes the phrase entered by the user, facial expression data, and voice data.

[0778] Input: User-entered phrases, facial expressions and voice data captured in real time

[0779] Output: Request data sent to the server

[0780] Step 4:

[0781] The server then puts the received phrases through an automated review process. The review uses natural language processing (NLP) technology to evaluate the appropriateness of the phrase. If the phrase contains inappropriate content, an error message is generated and returned to the terminal. If the phrase is appropriate, the process proceeds to the next step.

[0782] Input: Phrase sent to the server

[0783] Output: Evaluation result for suitability, error message or approval flag

[0784] Step 5:

[0785] The server sends the facial and voice data acquired in real time to the emotion recognition device, which then analyzes the data and recognizes the user's emotions (such as Microsoft Azure's Face API or Google Cloud's Vision AI). The recognized emotion data is then sent to the server.

[0786] Input: Real-time facial and voice data

[0787] Output: Recognized emotion data

[0788] Step 6:

[0789] Based on the phrases and recognized emotion data that have passed the screening process, the server uses a voice generation engine (OpenAI's GPT-3) to generate a custom voice message in the voice of the specified talent or character, taking into account emotional tone during the generation process.

[0790] Input: phrases that passed the screening, recognized emotion data

[0791] Output: Generated audio data

[0792] Step 7:

[0793] The server sends the generated audio data to the user terminal. In this step, the generated audio data and a playback link are sent to the user terminal.

[0794] Input: Generated audio data

[0795] Output: Audio data and playback link sent to the user's device

[0796] Step 8:

[0797] The user's device will then play the received audio data, allowing the user to listen to it. Additionally, the user will have the ability to share the audio message on social media or messaging apps.

[0798] Input: Audio data and playback link sent to the user's device

[0799] Output: Played audio data and a shareable link

[0800] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0801] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0802] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0803] [Third embodiment]

[0804] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0805] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0806] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0807] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0808] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0809] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0810] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0811] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0812] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0813] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0814] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0815] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0816] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. An embodiment of this system will be described.

[0817] overview

[0818] The system is primarily comprised of a user, a device, and a server. Users input and send the desired phrase using their device. The sent phrase is received by the server, and after its appropriateness is confirmed through an automatic review, it is generated as voice data in the voice of a designated talent or character using generative AI technology. The generated voice data is sent to the user's device, where the user can play or share it.

[0819] System Details

[0820] User Interface

[0821] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily input the desired phrase and send a request to generate a voice message.

[0822] Data transmission and reception

[0823] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character, typically as an HTTP POST request.

[0824] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0825] Automated Review Process

[0826] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0827] Speech Production Process

[0828] Phrases that pass the screening are input by the server into a voice generation means, where a generative AI model is used to generate voice data that mimics the voice of the talent or character selected by the user. A machine learning algorithm has been pre-trained for this generation.

[0829] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0830] Sending and storing audio data

[0831] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice message at any time later.

[0832] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0833] Share function

[0834] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0835] This allows users to easily and quickly enjoy customized messages in the voice of a specific celebrity or character. As described above, the system of the present invention provides users with high-quality personalized voice messages through a multi-stage process.

[0836] The processing flow will be explained below.

[0837] Step 1:

[0838] Users can open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, and a send button.

[0839] Step 2:

[0840] The user enters the desired phrase in the input field, for example, "Happy birthday, I'll keep on supporting you!"

[0841] Step 3:

[0842] The user selects a voice for a talent or character and clicks the send button, which sends the entered phrase and selection from the device to the server.

[0843] Step 4:

[0844] The device sends the user's input data to the server via an HTTP POST request, which includes the user's phrase and the selected talent or character information.

[0845] Step 5:

[0846] The server processes the received request and stores it in a database, which contains the user's phrases and selections.

[0847] Step 6:

[0848] The server puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrases.

[0849] Step 7:

[0850] If the automated review process detects an inappropriate phrase, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[0851] Step 8:

[0852] The server then feeds the selected phrases into a generative AI model, which uses machine learning algorithms to mimic the voice of the designated talent or character.

[0853] Step 9:

[0854] The generative AI model converts phrases into speech data in a specified voice. For example, the phrase "Happy birthday, I'll always support you!" is generated as speech data in the selected voice.

[0855] Step 10:

[0856] The generated voice data is then sent back to the server and stored in a database, allowing the user to access the voice data later.

[0857] Step 11:

[0858] The server generates a link or file URL to send the audio data to the user's device.

[0859] Step 12:

[0860] The server sends the generated link or file URL to the device, which receives it and displays it in its user interface.

[0861] Step 13:

[0862] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[0863] Step 14:

[0864] Users can share the generated voice message on social media or messaging applications, for example, using LINE or Twitter.

[0865] Example 1

[0866] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0867] With conventional technology, it was difficult for users to properly review and generate high-quality voice messages using the voices of specific celebrities or characters. Furthermore, there was a lack of a mechanism for evaluating the appropriateness of phrases or a way to efficiently share the generated voice data, which hindered user convenience.

[0868] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0869] In this invention, the server includes an input means for a user to input phrases, a transmission means for transmitting the input phrases to the server, a review means for automatically reviewing the transmitted phrases, a voice generation means for converting phrases that pass the review into specified voice, a storage means for saving the generated voice data in a database and generating a link or file URL for the saved voice data, a transmission means for sending the generated link or file URL to the user's terminal, and a playback means for the user to play or download the generated voice data. This makes it possible to efficiently create and share high-quality custom voice messages while ensuring the appropriateness of the phrases.

[0870] "User" means an individual who uses the System to generate custom voice messages using the voice of a particular talent or character.

[0871] "Input means" refers to the interface used by the user to input phrases, including a browser or dedicated application.

[0872] "Submission method" refers to the mechanism by which the entered phrase is sent to the server, typically as an HTTP POST request.

[0873] "Reviewing means" refers to a function for evaluating the appropriateness of submitted phrases. Specifically, it uses natural language processing (NLP) technology to automatically determine the appropriateness of phrases and whether they contain prohibited words.

[0874] "Voice generation means" refers to the technology used to convert the phrases that pass the screening process into a specified voice, using machine learning algorithms to imitate the voice of a specific talent or character.

[0875] "Storage Means" refers to the mechanism for storing the generated audio data in a database and generating a link or file URL to it.

[0876] "Playback means" refers to the function of sending the generated link or file URL to the user's device so that the user can play or download it.

[0877] "Database" means a system for storing transmitted phrases and generated speech data so that they can be accessed at a later time.

[0878] "Link" refers to the URL through which a user can access the generated audio data. By clicking on this link, the user can access the audio data.

[0879] "Machine learning algorithms" refer to programs trained to mimic the voices of specific talent or characters, also known as generative AI models.

[0880] "Natural Language Processing (NLP)" refers to techniques for assessing the appropriateness of phrases. It is used to detect profanity and evaluate the appropriateness of expressions.

[0881] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. The following describes an embodiment of this system.

[0882] System Configuration

[0883] This system mainly consists of users, terminals, and servers.

[0884] Hardware and software used

[0885] A terminal is a computing device (such as a personal computer, smartphone, or tablet) for user input, on which a browser or a dedicated application is installed.

[0886] The server has the computing resources to perform all functions such as receiving data, processing, generating audio, storing data in a database, generating links, etc. The server includes a web server, a database server, and an AI server on which the generative AI model runs.

[0887] A generative AI model is software that uses machine learning algorithms to mimic the voice of a specified talent or character.

[0888] Use software that uses natural language processing (NLP) techniques to determine the appropriateness of a phrase.

[0889] Processing flow and specific examples

[0890] Users can enter the phrase "Happy birthday, I'll keep cheering for you!" and select a specific talent or character. This can be done using the device's browser or a dedicated application.

[0891] The device sends the entered phrase and the selected talent or character information to the server as an HTTP POST request, typically using the HTTPS protocol.

[0892] The server processes the received request and stores it in a database, including the phrase entered by the user and the selected talent or character.

[0893] The server then automatically reviews the stored phrases using natural language processing (NLP) techniques, which evaluates whether the phrases are appropriate and whether they contain any banned words.

[0894] Phrases that pass the screening are input by the server into a voice generation means, which uses a generative AI model to generate voice data in the voice of the talent or character selected by the user. The generative AI model is pre-trained with a machine learning algorithm.

[0895] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[0896] The generated audio data is returned to the server and stored in a database. The server then generates a link or file URL for the saved audio data and sends it to the user's device.

[0897] The device displays the link or URL sent from the server in the user interface, and the user can click the link to play or download the audio data.

[0898] Finally, users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0899] This allows users to efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0900] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0901] Step 1:

[0902] The user accesses the interface by opening a browser or a dedicated application on their device. The user interface displays a phrase input field, a talent or character selection menu, and a send button. The user inputs the phrase "Happy birthday, I'll keep on cheering for you!" and selects a specific talent or character. The input is captured as keystrokes on the device.

[0903] Step 2:

[0904] The device creates an HTTP POST request based on the phrase entered by the user and the selected talent or character information. This request includes the text data entered by the user and the selected information. The created request is sent to the server using the HTTPS protocol. The input is the user's selected information, and the output is the sent HTTP request.

[0905] Step 3:

[0906] The server receives the incoming HTTP POST request and processes its contents. It extracts the phrase and talent or character information contained in the request and stores it in a database. Specifically, it executes an insert query to the database to create the corresponding record. The input is the HTTP request data, and the output is the updated database result.

[0907] Step 4:

[0908] The server then puts the stored phrases through an automated review process. The review process uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate. Here, the phrase is input into an NLP model to determine its appropriateness and the presence of prohibited words. For example, a phrase such as "happy birthday" would be reviewed. The input is the phrase read from the database, and the output is the review result.

[0909] Step 5:

[0910] Phrases that pass the screening are input by the server into a voice generation means. A generative AI model is used to generate voice data in the voice of a talent or character selected by the user. A machine learning algorithm can then be used to convert this phrase into the specified voice. The input is the phrase that passed the screening and the selected voice model, and the output is the generated voice data.

[0911] Step 6:

[0912] The generated audio data is returned to the server and stored in a database. The stored audio data includes the newly generated audio file and its link information. The input is the generated audio data, and the output is the updated database.

[0913] Step 7:

[0914] The server generates a link or file URL for the audio data stored in the database and sends it to the user's device. Specifically, it generates a link to the stored audio data as an HTTP response and sends it to the device. The input is the audio data record in the database, and the output is the HTTP response.

[0915] Step 8:

[0916] The terminal displays the links and URLs sent from the server in the user interface. The user can click the presented links to play or download the generated audio data. The input is the response data from the server, and the output is the updated result of the user interface.

[0917] Step 9:

[0918] Users share the generated audio message directly from their device to social media or messaging applications, for example by posting the generated audio link on social media platforms such as LINE or Twitter. The input is the generated link, and the output is the social media post.

[0919] In this way, users can efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[0920] (Application example 1)

[0921] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0922] Improving the user experience is a key challenge for modern food delivery services. In particular, there is a need for a fun way for users to track the status of their orders. Traditional notification methods are mainly text-based, which can be monotonous for users. Therefore, there is a need for a way to significantly improve the user experience by generating custom voice messages in the voices of specific celebrities or characters to notify delivery status.

[0923] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0924] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, message generation means for generating a specific status notification message using the generated voice data, and playback means for enabling the user to play or share the specific status notification message. This allows the user to receive delivery status notifications in the voice of a specific celebrity or character, and to enjoy checking the progress of their delivery.

[0925] A "user" is a person who can utilize the system to enter phrases to generate and play custom voice messages.

[0926] A "phrase" is text that a user enters to be generated as a voice message in the voice of a particular talent or character.

[0927] "Input means" refers to the interface and device that a user uses to input a phrase.

[0928] "Transmission means" is a communication means having the function of transmitting the input phrase to the server.

[0929] "Screening Measures" are the processes and functions for automatically screening submitted phrases and assessing their appropriateness.

[0930] "Voice generation means" refers to the technology and algorithms used to convert phrases that pass screening into the voice of a specific talent or character.

[0931] "Audio data" refers to a digital audio file generated by an audio generating means.

[0932] The "message generator" is a system for constructing a specific status notification message using the generated voice data.

[0933] "Playback means" refers to the device and interface that allows a user to play or share a generated voice message.

[0934] This invention is a system that allows users to receive specific status notification messages in the voice of a specific celebrity or character. This system is mainly composed of a user, a terminal, a server, and a database.

[0935] User Interface

[0936] Users access the system through a smartphone application, which includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily enter the desired phrase and send a request to the server to generate a voice message.

[0937] Data transmission and reception

[0938] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character. The request is typically sent as an HTTP POST request. The server receives this request and first stores it in a database. The stored data contains the information needed for subsequent processing.

[0939] Automated Review Process

[0940] The server includes a reviewing means for automatically reviewing received phrases. The reviewing means uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate and whether it contains prohibited words. If the phrase contains inappropriate language, an error message is generated.

[0941] Speech Production Process

[0942] Phrases that pass the screening are input into the voice generation means by the server. A generative AI model is used to generate voice data in the voice of the talent or character selected by the user. A machine learning algorithm is pre-trained for generation. For example, if a user sends the phrase "Thank you for your order! Your food is being prepared," the phrase is generated in the voice of a specific talent.

[0943] Message Generation Process

[0944] The generated voice data is configured as a specific situation notification message, such as "Your food is ready! We'll deliver it to you shortly."

[0945] Sending and storing audio data

[0946] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice message at any time later. The server generates a link or file URL to send the voice data to the user's device. This link or URL is sent to the device and displayed in the user interface. The user can click it to play or download the voice data.

[0947] Share function

[0948] Users can share the generated voice messages directly from their devices to social media or messaging applications, such as LINE or Twitter, with friends and family. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters.

[0949] Prompt Sentence Examples

[0950] For example, use the following prompt:

[0951] "Thank you for your order! Your food is ready."

[0952] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0953] Step 1:

[0954] The user inputs a phrase through the application. They enter the phrase in the input field, select a talent or character, and press the send button to complete the input. The notification message entered here is the user's desired message.

[0955] Step 2:

[0956] The device sends a request to the server containing the entered phrase and information about the selected talent or character. This request is sent as an HTTP POST request, and the server receives it. The received data includes the phrase and character information.

[0957] Step 3:

[0958] The server stores the received phrase in a database. The stored data includes information such as the user ID, phrase, selected characters, and timestamp. This preserves the information necessary for subsequent processing.

[0959] Step 4:

[0960] The server then puts the stored phrases through an automated review process. It uses natural language processing (NLP) techniques to determine whether the phrase is appropriate and whether it contains any forbidden words. If it does contain any inappropriate language, it generates an error message and sends it back to the user. This step takes the phrase as input and the review result as output.

[0961] Step 5:

[0962] Phrases that pass the screening are input into a voice generation means. The server uses a machine learning algorithm to convert the phrase into voice data in the voice of the designated talent or character. A generative AI model is responsible for this process. Here, the input is the passed phrase and character information, and the output is voice data.

[0963] Step 6:

[0964] The generated audio data is then sent back to the server and stored in a database. The stored data includes the URL of the audio file and metadata. The user receives a link or file URL as the information needed to access this data.

[0965] Step 7:

[0966] The server uses the generated voice data to generate a specific status notification message and sends it to the user's device. An example of a notification message is "Your food is ready! We'll deliver it to you shortly." The input here is the voice data and status information, and the output is the generated notification message.

[0967] Step 8:

[0968] The user receives a notification message on their device and plays or shares it. By pressing the play button on the application interface, the audio message is played. By pressing the share button, the message can be shared on social media such as LINE or Twitter. Here, the notification message is the input, and audio playback or sharing is the output.

[0969] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0970] The present invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. The following describes in detail the embodiments of the invention.

[0971] overview

[0972] The system consists of a user, a device, a server, and an emotion engine. Users use their device to input and send a desired phrase. The sent phrase is received by the server, and after its appropriateness is confirmed through automatic review, it is generated as voice data using AI technology in a voice tone that corresponds to the user's emotion recognized by the emotion engine. The generated voice data is sent to the user's device, where the user can play or share it.

[0973] System Details

[0974] User Interface

[0975] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone permission screen. This allows users to easily input the desired phrase and send a request to generate a voice message.

[0976] Data transmission and reception

[0977] The device sends a request to the server, typically as an HTTP POST request, containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine.

[0978] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[0979] Automated Review Process

[0980] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[0981] Emotion Recognition Process

[0982] Phrases that pass the screening process are sent to an emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[0983] Speech Production Process

[0984] The emotion information recognized by the emotion engine is sent to a server and linked to a voice generation means. Phrases and emotion information that pass the screening are input into a generative AI model. This AI model uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[0985] For example, if a user sends the phrase "Happy birthday, I'll keep on supporting you!" and the emotion engine recognizes a positive emotion, it will generate a voice with a bright tone that reflects that emotion.

[0986] Sending and storing audio data

[0987] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[0988] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[0989] Share function

[0990] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[0991] This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[0992] The processing flow will be explained below.

[0993] Step 1:

[0994] Users open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, a send button, and a camera and microphone permission screen.

[0995] Step 2:

[0996] The user inputs a desired phrase into the input field. For example, the user inputs a phrase such as "Happy birthday, I'll continue to support you!"

[0997] Step 3:

[0998] The user selects the voice of the talent or character, grants permission to access the camera and microphone, and then clicks the send button, which sends the entered phrase, selection information, and image and audio data from the device to the server.

[0999] Step 4:

[1000] The device sends input data and image and audio data required for emotion recognition to the server via an HTTP POST request. This request includes the user's phrase, information about the selected talent or character, and data used by the emotion engine.

[1001] Step 5:

[1002] The server processes the received data and first stores it in a database, which contains the information needed for subsequent processing.

[1003] Step 6:

[1004] The server then puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but if it contains inappropriate language, an error message is generated.

[1005] Step 7:

[1006] If an inappropriate phrase is detected during the automated review process, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[1007] Step 8:

[1008] Phrases that pass the screening are sent by the server to the emotion engine, which estimates the user's emotions from the images and audio data captured by the user's camera and microphone.

[1009] Step 9:

[1010] The emotion engine uses facial recognition and voice analysis technologies to estimate emotions from the user's facial expressions and tone of voice. For example, a smile is recognized as a positive emotion, while tears are recognized as a negative emotion.

[1011] Step 10:

[1012] The recognized emotion information is sent to a server and integrated with the speech generation means, which provides the emotion information to the speech generation means so that phrases are generated with a tone and inflection appropriate to the emotion.

[1013] Step 11:

[1014] The phrases and emotional information that pass the screening process are then fed into a generative AI model, which uses machine learning algorithms to generate audio data that mimics the voice of the designated talent or character, adding tone and inflection that corresponds to the recognized emotion.

[1015] Step 12:

[1016] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[1017] Step 13:

[1018] The server generates a link or file URL to send the audio data to the user's device, which then displays the link or URL in a user interface.

[1019] Step 14:

[1020] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[1021] Step 15:

[1022] Users can share the generated voice message directly from their device to social media or messaging applications, for example, using LINE or Twitter.

[1023] Example 2

[1024] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1025] Conventional voice message generation systems have difficulty generating voice messages that reflect the user's emotions, preventing them from providing a personalized experience. Furthermore, the process for evaluating the appropriateness of phrases is insufficient, which can result in inappropriate phrases being generated. Furthermore, when imitating the voices of specific celebrities or characters, natural tones and intonations are often lacking.

[1026] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1027] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for recognizing emotions, and means for adjusting the tone and intonation of the voice in accordance with the emotion recognized by the emotion recognition means. This makes it possible to generate a personalized voice message that reflects the user's emotions and provide a more natural and appropriate voice message.

[1028] An "input means" is a device or software that provides an interface for a user to input a phrase.

[1029] "Transmission means" refers to a device or software that has the function of transmitting phrases and information entered by the user to the server.

[1030] "Screening means" refers to a device or software that has the function of automatically screening transmitted phrases and evaluating their appropriateness and the presence of prohibited words.

[1031] "Speech generation means" refers to a device or software that converts phrases that have passed the screening into a specified voice and generates voice data.

[1032] "Emotion recognition means" refers to devices or software that use algorithms to recognize emotions from the user's facial expressions and voice via the user's camera or microphone.

[1033] The "tone and intonation adjusting means" refers to a device or software that has the function of adjusting the tone and intonation of the generated voice in accordance with the user's emotion recognized by the emotion recognition means.

[1034] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. Components of the system include a user, a terminal, a server, and the emotion engine.

[1035] A user inputs a desired phrase using their own device and sends it. A browser or a dedicated application is installed on the device, which functions as a user interface. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone access permission screen.

[1036] The device generates a request containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine, and sends this request to the server as an HTTP POST request.

[1037] The server receives this request and stores it in a database. The stored data contains information necessary for subsequent processing. The server also subjects the received phrases to an automated review process. This review process uses natural language processing (NLP) techniques to evaluate the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but an error message is generated if it contains inappropriate language.

[1038] Phrases that pass the screening process are sent to the emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[1039] The emotion information recognized by the emotion engine is sent to a server, which then connects to a voice generator. The server then inputs the selected phrases and emotion information into a generative AI model, which uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[1040] As an example of a specific prompt, let's say the user inputs the phrase "Happy birthday! Stay healthy!" in Talent A's voice while smiling. The emotion engine recognizes the user's smile as a positive emotion, and generates a bright-toned voice that reflects that emotion. Here's a concrete example of how this prompt can be input into the generative AI model:

[1041] "Generate the phrase 'Happy birthday! Stay healthy!' in Talent A's voice based on the positive emotions expressed by smiling users."

[1042] The generated audio data is returned to the server and stored in a database. This storage allows the user to access the audio data at any time later. The server generates a link or file URL to send the audio data to the user's device, and this link or URL is sent to the device. The user can play or download the audio data by clicking the link or URL displayed in the user interface.

[1043] Users can also share the generated voice messages directly from their devices on social media and messaging applications. For example, they can share messages with friends and family using LINE or Twitter. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[1044] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1045] Step 1: User enters phrase

[1046] The user accesses the system via the device's browser or a dedicated application. The user selects a phrase to enter in the phrase input field displayed in the user interface and a talent or character from a selection menu. Specifically, the user enters the phrase "Happy birthday! Stay healthy!" and selects Talent A. The entered data is stored internally on the device as a generated prompt.

[1047] input:

[1048] Phrase: "Happy birthday! Stay healthy!"

[1049] Talent or Character: Talent A

[1050] output:

[1051] Generated prompt data

[1052] Step 2: Generate and send a request

[1053] The device generates request data using the entered phrase and information about the selected talent or character, and sends it to the server as an HTTP POST request.

[1054] input:

[1055] Prompt Data

[1056] Data processing and calculation:

[1057] Generating an HTTP POST Request

[1058] output:

[1059] Sending requests to the server

[1060] Step 3: Data received and stored by the server

[1061] The server receives the request sent from the device and stores it in a database, including phrases and talent or character information.

[1062] input:

[1063] Send Request

[1064] Data processing and calculation:

[1065] Saving to a database

[1066] output:

[1067] Phrases and talent information stored in the database

[1068] Step 4: Automated Review Process

[1069] The server automatically reviews stored phrases for appropriateness using natural language processing (NLP) techniques: for example, "Happy Birthday" is deemed appropriate, while inappropriate phrases generate an error message.

[1070] input:

[1071] Phrases in the database

[1072] Data processing and calculation:

[1073] Appropriateness assessment using natural language processing

[1074] output:

[1075] Review result (appropriate or inappropriate)

[1076] Step 5: Emotion Recognition Process

[1077] Phrases deemed appropriate are sent to the emotion engine, which recognizes emotions from facial expressions and tone of voice via the user's camera or microphone. For example, if the user is smiling, it is recognized as a positive emotion.

[1078] input:

[1079] Phrase (reviewed)

[1080] User facial expression data

[1081] User's voice tone

[1082] Data processing and calculation:

[1083] Applying emotion recognition algorithms

[1084] output:

[1085] Recognized emotional information

[1086] Step 6: The speech generation process

[1087] Based on the emotional information from the emotion engine and the screened phrases, the server uses a generative AI model to generate audio in the voice of the designated talent or character, with tone and inflection adjusted according to the recognized emotion.

[1088] input:

[1089] phrase

[1090] Recognized emotional information

[1091] Data processing and calculation:

[1092] Speech generation using generative AI models

[1093] output:

[1094] Generated audio data

[1095] Step 7: Send and save the audio data

[1096] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice data at any time later. The server generates a link or file URL to send the voice data to the user's device and sends this link to the user.

[1097] input:

[1098] Generated audio data

[1099] Data processing and calculation:

[1100] Saving to a database

[1101] Generate a download link or URL

[1102] output:

[1103] Links or URLs sent to users

[1104] Step 8: Sharing Function

[1105] Users receive the generated audio message and can click a link or URL to play or download the audio data, or share the audio message with friends and family via social media or messaging applications.

[1106] input:

[1107] Links or URLs sent to users

[1108] Data processing and calculation:

[1109] Download or play audio data

[1110] Share audio messages

[1111] output:

[1112] Playing audio data

[1113] Share audio messages

[1114] (Application example 2)

[1115] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1116] In modern advertising, sending unilateral messages without considering user emotions cannot be expected to be effective. Furthermore, uniform voice messages have difficulty attracting user attention, and a lack of personalization leads to reduced engagement with advertising. Therefore, there is a need for a system that can grasp a user's emotional state in real time and generate and deliver personalized advertising messages with appropriate voice tones.

[1117] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for the user to input a phrase, a transmission means for transmitting the input phrase to the server, an examination means for automatically examining the transmitted phrase, an emotion recognition means for recognizing an emotion, a voice generation means for converting the phrase into a specified voice with an audio tone corresponding to the emotion recognized by the emotion recognition means, and a transmission means for transmitting the generated voice data to the user's terminal and playing it back. This makes it possible to generate and deliver a personalized voice message that takes the user's emotions into consideration.

[1118] "User" means a person or entity that utilizes the system to input phrases and generate voice messages.

[1119] "Input means" refers to a device or software function that allows a user to input a desired phrase.

[1120] "Transmission means" refers to a device or software function for transmitting the input phrase to the server.

[1121] A "screening tool" is a process or algorithm that automatically determines the appropriateness of a submitted phrase.

[1122] "Emotion recognition means" refers to a device or software function that reads and analyzes a user's emotions through a camera or microphone.

[1123] A "voice generation means" is a device or software function that converts phrases into voice data based on a specified voice or emotional tone.

[1124] The "server" is a central processing unit that processes various data and transmits audio data and control information to client terminals.

[1125] A "terminal" is a device (such as a smartphone or PC) that a user directly operates and interacts with the system.

[1126] A "phrase" is a series of words or sentences that a user wishes to convert into a voice message.

[1127] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. This system incorporates an emotion engine that recognizes the user's emotions, enabling the creation of more personalized voice messages.

[1128] composition

[1129] This system consists of a user terminal, a server, an emotion recognition engine, and a voice generation engine. The user terminal is assumed to be a smartphone, but other devices can also be used. The server acts as a central processing unit and manages the emotion recognition and voice generation processes.

[1130] Hardware / Software used

[1131] Hardware:

[1132] Smartphone (iOS, Android)

[1133] software:

[1134] Microsoft Azure's Face API or Google Cloud's Vision AI (emotion recognition)

[1135] OpenAI's GPT-3 (Speech Generation)

[1136] HTTP server (data transmission and reception)

[1137] Operation method and configuration

[1138] 1. User Interface:

[1139] The user launches the app and is asked for camera and microphone permissions, which initiates the emotion recognition process. The user then enters the desired message in the phrase input field and presses send.

[1140] 2. Emotion recognition:

[1141] Data is collected in real time from the camera and microphone on the user's device and sent to the emotion recognition API, where the user's facial expressions and tone of voice are analyzed to obtain emotional information.

[1142] 3. Data transmission:

[1143] The acquired emotion information and the input phrase are sent to the server using an HTTP POST request.

[1144] 4. Automated Review:

[1145] The server receives the submitted phrase and evaluates its appropriateness using natural language processing (NLP) techniques. If it contains inappropriate language, it returns an error message; if it is an appropriate phrase, it proceeds to the next step.

[1146] 5. Speech generation:

[1147] Based on the phrases and emotional information that pass the screening process, a speech generation engine (OpenAI's GPT-3) is used to generate a custom voice message in the voice of the designated talent or character, with the tone and intonation adjusted according to the recognized emotion.

[1148] 6. Data transmission and playback:

[1149] The generated audio data is then sent back to the user's device via the server, where it can be played within the application, and a sharing link is also generated, allowing the user to share the audio message on social media or via messaging apps.

[1150] Specific examples

[1151] For example, if a user is in a good mood and types the phrase "Thank you for your hard work today!", the emotion recognition engine will recognize the positive emotion and generate a message in a bright tone such as "Thank you for your hard work today!"

[1152] Example prompt sentence:

[1153] "Please generate an advertising message in the voice of celebrity X to be played when the user has a happy expression. The message should introduce product Z in a cheerful tone."

[1154] In this way, it becomes possible to generate and deliver personalized voice messages that take into account the user's emotions.

[1155] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1156] Step 1:

[1157] The user launches the application and grants permission to use the camera and microphone. Once this is complete, the system is ready to collect facial and voice data in real time.

[1158] Input: App launch, camera and microphone permissions

[1159] Output: Real-time facial expression and voice data acquisition status

[1160] Step 2:

[1161] The user enters the phrase using an input device and presses the send button, at which point the phrase is saved on the device and ready to be sent in the next step.

[1162] Input: A phrase entered by the user

[1163] Output: The phrase you entered is saved to your device.

[1164] Step 3:

[1165] The user device sends the phrase entered and the facial expression and voice data acquired in real time to the server as an HTTP POST request. The transmitted data includes the phrase entered by the user, facial expression data, and voice data.

[1166] Input: User-entered phrases, facial expressions and voice data captured in real time

[1167] Output: Request data sent to the server

[1168] Step 4:

[1169] The server then puts the received phrases through an automated review process. The review uses natural language processing (NLP) technology to evaluate the appropriateness of the phrase. If the phrase contains inappropriate content, an error message is generated and returned to the terminal. If the phrase is appropriate, the process proceeds to the next step.

[1170] Input: Phrase sent to the server

[1171] Output: Evaluation result for suitability, error message or approval flag

[1172] Step 5:

[1173] The server sends the facial and voice data acquired in real time to the emotion recognition device, which then analyzes the data and recognizes the user's emotions (such as Microsoft Azure's Face API or Google Cloud's Vision AI). The recognized emotion data is then sent to the server.

[1174] Input: Real-time facial and voice data

[1175] Output: Recognized emotion data

[1176] Step 6:

[1177] Based on the phrases and recognized emotion data that have passed the screening process, the server uses a voice generation engine (OpenAI's GPT-3) to generate a custom voice message in the voice of the specified talent or character, taking into account emotional tone during the generation process.

[1178] Input: phrases that passed the screening, recognized emotion data

[1179] Output: Generated audio data

[1180] Step 7:

[1181] The server sends the generated audio data to the user terminal. In this step, the generated audio data and a playback link are sent to the user terminal.

[1182] Input: Generated audio data

[1183] Output: Audio data and playback link sent to the user's device

[1184] Step 8:

[1185] The user's device will then play the received audio data, allowing the user to listen to it. Additionally, the user will have the ability to share the audio message on social media or messaging apps.

[1186] Input: Audio data and playback link sent to the user's device

[1187] Output: Played audio data and a shareable link

[1188] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1191] [Fourth embodiment]

[1192] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1193] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1194] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1195] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1196] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1198] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1199] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1200] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1201] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1202] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1203] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1204] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1205] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. An embodiment of this system will be described.

[1206] overview

[1207] The system is primarily comprised of a user, a device, and a server. Users input and send the desired phrase using their device. The sent phrase is received by the server, and after its appropriateness is confirmed through an automatic review, it is generated as voice data in the voice of a designated talent or character using generative AI technology. The generated voice data is sent to the user's device, where the user can play or share it.

[1208] System Details

[1209] User Interface

[1210] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily input the desired phrase and send a request to generate a voice message.

[1211] Data transmission and reception

[1212] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character, typically as an HTTP POST request.

[1213] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[1214] Automated Review Process

[1215] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[1216] Speech Production Process

[1217] Phrases that pass the screening are input by the server into a voice generation means, where a generative AI model is used to generate voice data that mimics the voice of the talent or character selected by the user. A machine learning algorithm has been pre-trained for this generation.

[1218] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[1219] Sending and storing audio data

[1220] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice message at any time later.

[1221] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[1222] Share function

[1223] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[1224] This allows users to easily and quickly enjoy customized messages in the voice of a specific celebrity or character. As described above, the system of the present invention provides users with high-quality personalized voice messages through a multi-stage process.

[1225] The processing flow will be explained below.

[1226] Step 1:

[1227] Users can open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, and a send button.

[1228] Step 2:

[1229] The user enters the desired phrase in the input field, for example, "Happy birthday, I'll keep on supporting you!"

[1230] Step 3:

[1231] The user selects a voice for a talent or character and clicks the send button, which sends the entered phrase and selection from the device to the server.

[1232] Step 4:

[1233] The device sends the user's input data to the server via an HTTP POST request, which includes the user's phrase and the selected talent or character information.

[1234] Step 5:

[1235] The server processes the received request and stores it in a database, which contains the user's phrases and selections.

[1236] Step 6:

[1237] The server puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrases.

[1238] Step 7:

[1239] If the automated review process detects an inappropriate phrase, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[1240] Step 8:

[1241] The server then feeds the selected phrases into a generative AI model, which uses machine learning algorithms to mimic the voice of the designated talent or character.

[1242] Step 9:

[1243] The generative AI model converts phrases into speech data in a specified voice. For example, the phrase "Happy birthday, I'll always support you!" is generated as speech data in the selected voice.

[1244] Step 10:

[1245] The generated voice data is then sent back to the server and stored in a database, allowing the user to access the voice data later.

[1246] Step 11:

[1247] The server generates a link or file URL to send the audio data to the user's device.

[1248] Step 12:

[1249] The server sends the generated link or file URL to the device, which receives it and displays it in its user interface.

[1250] Step 13:

[1251] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[1252] Step 14:

[1253] Users can share the generated voice message on social media or messaging applications, for example, using LINE or Twitter.

[1254] Example 1

[1255] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1256] With conventional technology, it was difficult for users to properly review and generate high-quality voice messages using the voices of specific celebrities or characters. Furthermore, there was a lack of a mechanism for evaluating the appropriateness of phrases or a way to efficiently share the generated voice data, which hindered user convenience.

[1257] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1258] In this invention, the server includes an input means for a user to input phrases, a transmission means for transmitting the input phrases to the server, a review means for automatically reviewing the transmitted phrases, a voice generation means for converting phrases that pass the review into specified voice, a storage means for saving the generated voice data in a database and generating a link or file URL for the saved voice data, a transmission means for sending the generated link or file URL to the user's terminal, and a playback means for the user to play or download the generated voice data. This makes it possible to efficiently create and share high-quality custom voice messages while ensuring the appropriateness of the phrases.

[1259] "User" means an individual who uses the System to generate custom voice messages using the voice of a particular talent or character.

[1260] "Input means" refers to the interface used by the user to input phrases, including a browser or dedicated application.

[1261] "Submission method" refers to the mechanism by which the entered phrase is sent to the server, typically as an HTTP POST request.

[1262] "Reviewing means" refers to a function for evaluating the appropriateness of submitted phrases. Specifically, it uses natural language processing (NLP) technology to automatically determine the appropriateness of phrases and whether they contain prohibited words.

[1263] "Voice generation means" refers to the technology used to convert the phrases that pass the screening process into a specified voice, using machine learning algorithms to imitate the voice of a specific talent or character.

[1264] "Storage Means" refers to the mechanism for storing the generated audio data in a database and generating a link or file URL to it.

[1265] "Playback means" refers to the function of sending the generated link or file URL to the user's device so that the user can play or download it.

[1266] "Database" means a system for storing transmitted phrases and generated speech data so that they can be accessed at a later time.

[1267] "Link" refers to the URL through which a user can access the generated audio data. By clicking on this link, the user can access the audio data.

[1268] "Machine learning algorithms" refer to programs trained to mimic the voices of specific talent or characters, also known as generative AI models.

[1269] "Natural Language Processing (NLP)" refers to techniques for assessing the appropriateness of phrases. It is used to detect profanity and evaluate the appropriateness of expressions.

[1270] The present invention relates to a system that allows a user to create and use a custom voice message using the voice of a specific celebrity or character. The following describes an embodiment of this system.

[1271] System Configuration

[1272] This system mainly consists of users, terminals, and servers.

[1273] Hardware and software used

[1274] A terminal is a computing device (such as a personal computer, smartphone, or tablet) for user input, on which a browser or a dedicated application is installed.

[1275] The server has the computing resources to perform all functions such as receiving data, processing, generating audio, storing data in a database, generating links, etc. The server includes a web server, a database server, and an AI server on which the generative AI model runs.

[1276] A generative AI model is software that uses machine learning algorithms to mimic the voice of a specified talent or character.

[1277] Use software that uses natural language processing (NLP) techniques to determine the appropriateness of a phrase.

[1278] Processing flow and specific examples

[1279] Users can enter the phrase "Happy birthday, I'll keep cheering for you!" and select a specific talent or character. This can be done using the device's browser or a dedicated application.

[1280] The device sends the entered phrase and the selected talent or character information to the server as an HTTP POST request, typically using the HTTPS protocol.

[1281] The server processes the received request and stores it in a database, including the phrase entered by the user and the selected talent or character.

[1282] The server then automatically reviews the stored phrases using natural language processing (NLP) techniques, which evaluates whether the phrases are appropriate and whether they contain any banned words.

[1283] Phrases that pass the screening are input by the server into a voice generation means, which uses a generative AI model to generate voice data in the voice of the talent or character selected by the user. The generative AI model is pre-trained with a machine learning algorithm.

[1284] For example, if a user sends the phrase "Happy birthday, I'll continue to support you!" and selects the voice of a specific idol, the message will be generated in that idol's voice.

[1285] The generated audio data is returned to the server and stored in a database. The server then generates a link or file URL for the saved audio data and sends it to the user's device.

[1286] The device displays the link or URL sent from the server in the user interface, and the user can click the link to play or download the audio data.

[1287] Finally, users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[1288] This allows users to efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[1289] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1290] Step 1:

[1291] The user accesses the interface by opening a browser or a dedicated application on their device. The user interface displays a phrase input field, a talent or character selection menu, and a send button. The user inputs the phrase "Happy birthday, I'll keep on cheering for you!" and selects a specific talent or character. The input is captured as keystrokes on the device.

[1292] Step 2:

[1293] The device creates an HTTP POST request based on the phrase entered by the user and the selected talent or character information. This request includes the text data entered by the user and the selected information. The created request is sent to the server using the HTTPS protocol. The input is the user's selected information, and the output is the sent HTTP request.

[1294] Step 3:

[1295] The server receives the incoming HTTP POST request and processes its contents. It extracts the phrase and talent or character information contained in the request and stores it in a database. Specifically, it executes an insert query to the database to create the corresponding record. The input is the HTTP request data, and the output is the updated database result.

[1296] Step 4:

[1297] The server then puts the stored phrases through an automated review process. The review process uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate. Here, the phrase is input into an NLP model to determine its appropriateness and the presence of prohibited words. For example, a phrase such as "happy birthday" would be reviewed. The input is the phrase read from the database, and the output is the review result.

[1298] Step 5:

[1299] Phrases that pass the screening are input by the server into a voice generation means. A generative AI model is used to generate voice data in the voice of a talent or character selected by the user. A machine learning algorithm can then be used to convert this phrase into the specified voice. The input is the phrase that passed the screening and the selected voice model, and the output is the generated voice data.

[1300] Step 6:

[1301] The generated audio data is returned to the server and stored in a database. The stored audio data includes the newly generated audio file and its link information. The input is the generated audio data, and the output is the updated database.

[1302] Step 7:

[1303] The server generates a link or file URL for the audio data stored in the database and sends it to the user's device. Specifically, it generates a link to the stored audio data as an HTTP response and sends it to the device. The input is the audio data record in the database, and the output is the HTTP response.

[1304] Step 8:

[1305] The terminal displays the links and URLs sent from the server in the user interface. The user can click the presented links to play or download the generated audio data. The input is the response data from the server, and the output is the updated result of the user interface.

[1306] Step 9:

[1307] Users share the generated audio message directly from their device to social media or messaging applications, for example by posting the generated audio link on social media platforms such as LINE or Twitter. The input is the generated link, and the output is the social media post.

[1308] In this way, users can efficiently create and share high-quality custom voice messages while ensuring phrasing is appropriate.

[1309] (Application example 1)

[1310] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1311] Improving the user experience is a key challenge for modern food delivery services. In particular, there is a need for a fun way for users to track the status of their orders. Traditional notification methods are mainly text-based, which can be monotonous for users. Therefore, there is a need for a way to significantly improve the user experience by generating custom voice messages in the voices of specific celebrities or characters to notify delivery status.

[1312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1313] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, message generation means for generating a specific status notification message using the generated voice data, and playback means for enabling the user to play or share the specific status notification message. This allows the user to receive delivery status notifications in the voice of a specific celebrity or character, and to enjoy checking the progress of their delivery.

[1314] A "user" is a person who can utilize the system to enter phrases to generate and play custom voice messages.

[1315] A "phrase" is text that a user enters to be generated as a voice message in the voice of a particular talent or character.

[1316] "Input means" refers to the interface and device that a user uses to input a phrase.

[1317] "Transmission means" is a communication means having the function of transmitting the input phrase to the server.

[1318] "Screening Measures" are the processes and functions for automatically screening submitted phrases and assessing their appropriateness.

[1319] "Voice generation means" refers to the technology and algorithms used to convert phrases that pass screening into the voice of a specific talent or character.

[1320] "Audio data" refers to a digital audio file generated by an audio generating means.

[1321] The "message generator" is a system for constructing a specific status notification message using the generated voice data.

[1322] "Playback means" refers to the device and interface that allows a user to play or share a generated voice message.

[1323] This invention is a system that allows users to receive specific status notification messages in the voice of a specific celebrity or character. This system is mainly composed of a user, a terminal, a server, and a database.

[1324] User Interface

[1325] Users access the system through a smartphone application, which includes a phrase input field, a talent or character selection menu, and a send button, allowing users to easily enter the desired phrase and send a request to the server to generate a voice message.

[1326] Data transmission and reception

[1327] The device sends a request to the server containing the phrase entered by the user and information about the selected talent or character. The request is typically sent as an HTTP POST request. The server receives this request and first stores it in a database. The stored data contains the information needed for subsequent processing.

[1328] Automated Review Process

[1329] The server includes a reviewing means for automatically reviewing received phrases. The reviewing means uses natural language processing (NLP) techniques to evaluate whether the phrase is appropriate and whether it contains prohibited words. If the phrase contains inappropriate language, an error message is generated.

[1330] Speech Production Process

[1331] Phrases that pass the screening are input into the voice generation means by the server. A generative AI model is used to generate voice data in the voice of the talent or character selected by the user. A machine learning algorithm is pre-trained for generation. For example, if a user sends the phrase "Thank you for your order! Your food is being prepared," the phrase is generated in the voice of a specific talent.

[1332] Message Generation Process

[1333] The generated voice data is configured as a specific situation notification message, such as "Your food is ready! We'll deliver it to you shortly."

[1334] Sending and storing audio data

[1335] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice message at any time later. The server generates a link or file URL to send the voice data to the user's device. This link or URL is sent to the device and displayed in the user interface. The user can click it to play or download the voice data.

[1336] Share function

[1337] Users can share the generated voice messages directly from their devices to social media or messaging applications, such as LINE or Twitter, with friends and family. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters.

[1338] Prompt Sentence Examples

[1339] For example, use the following prompt:

[1340] "Thank you for your order! Your food is ready."

[1341] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1342] Step 1:

[1343] The user inputs a phrase through the application. They enter the phrase in the input field, select a talent or character, and press the send button to complete the input. The notification message entered here is the user's desired message.

[1344] Step 2:

[1345] The device sends a request to the server containing the entered phrase and information about the selected talent or character. This request is sent as an HTTP POST request, and the server receives it. The received data includes the phrase and character information.

[1346] Step 3:

[1347] The server stores the received phrase in a database. The stored data includes information such as the user ID, phrase, selected characters, and timestamp. This preserves the information necessary for subsequent processing.

[1348] Step 4:

[1349] The server then puts the stored phrases through an automated review process. It uses natural language processing (NLP) techniques to determine whether the phrase is appropriate and whether it contains any forbidden words. If it does contain any inappropriate language, it generates an error message and sends it back to the user. This step takes the phrase as input and the review result as output.

[1350] Step 5:

[1351] Phrases that pass the screening are input into a voice generation means. The server uses a machine learning algorithm to convert the phrase into voice data in the voice of the designated talent or character. A generative AI model is responsible for this process. Here, the input is the passed phrase and character information, and the output is voice data.

[1352] Step 6:

[1353] The generated audio data is then sent back to the server and stored in a database. The stored data includes the URL of the audio file and metadata. The user receives a link or file URL as the information needed to access this data.

[1354] Step 7:

[1355] The server uses the generated voice data to generate a specific status notification message and sends it to the user's device. An example of a notification message is "Your food is ready! We'll deliver it to you shortly." The input here is the voice data and status information, and the output is the generated notification message.

[1356] Step 8:

[1357] The user receives a notification message on their device and plays or shares it. By pressing the play button on the application interface, the audio message is played. By pressing the share button, the message can be shared on social media such as LINE or Twitter. Here, the notification message is the input, and audio playback or sharing is the output.

[1358] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1359] The present invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. The following describes in detail the embodiments of the invention.

[1360] overview

[1361] The system consists of a user, a device, a server, and an emotion engine. Users use their device to input and send a desired phrase. The sent phrase is received by the server, and after its appropriateness is confirmed through automatic review, it is generated as voice data using AI technology in a voice tone that corresponds to the user's emotion recognized by the emotion engine. The generated voice data is sent to the user's device, where the user can play or share it.

[1362] System Details

[1363] User Interface

[1364] Users access the system through a browser or a dedicated application. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone permission screen. This allows users to easily input the desired phrase and send a request to generate a voice message.

[1365] Data transmission and reception

[1366] The device sends a request to the server, typically as an HTTP POST request, containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine.

[1367] The server receives this request and first stores it in a database, which contains the information needed for subsequent processing.

[1368] Automated Review Process

[1369] The server then submits the received phrases to an automated review process that uses natural language processing (NLP) techniques to assess whether the phrase is appropriate and whether it contains any profanity. For example, the phrase "Happy Birthday" is considered appropriate, but if it contains profanity an error message is generated.

[1370] Emotion Recognition Process

[1371] Phrases that pass the screening process are sent to an emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[1372] Speech Production Process

[1373] The emotion information recognized by the emotion engine is sent to a server and linked to a voice generation means. Phrases and emotion information that pass the screening are input into a generative AI model. This AI model uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[1374] For example, if a user sends the phrase "Happy birthday, I'll keep on supporting you!" and the emotion engine recognizes a positive emotion, it will generate a voice with a bright tone that reflects that emotion.

[1375] Sending and storing audio data

[1376] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[1377] The server generates a link or file URL to send the audio data to the user's device. This link or URL is sent to the device and displayed in a user interface, where the user can click to play or download the audio data.

[1378] Share function

[1379] Users can share the generated voice messages directly from their device to social media or messaging applications, for example, using LINE or Twitter to share the messages they generate with friends and family.

[1380] This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[1381] The processing flow will be explained below.

[1382] Step 1:

[1383] Users open a custom voice message creation screen through a browser or dedicated application. The user interface includes a phrase entry field, a talent or character selection menu, a send button, and a camera and microphone permission screen.

[1384] Step 2:

[1385] The user inputs a desired phrase into the input field. For example, the user inputs a phrase such as "Happy birthday, I'll continue to support you!"

[1386] Step 3:

[1387] The user selects the voice of the talent or character, grants permission to access the camera and microphone, and then clicks the send button, which sends the entered phrase, selection information, and image and audio data from the device to the server.

[1388] Step 4:

[1389] The device sends input data and image and audio data required for emotion recognition to the server via an HTTP POST request. This request includes the user's phrase, information about the selected talent or character, and data used by the emotion engine.

[1390] Step 5:

[1391] The server processes the received data and first stores it in a database, which contains the information needed for subsequent processing.

[1392] Step 6:

[1393] The server then puts the stored phrases through an automated review process, using natural language processing (NLP) algorithms to assess the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but if it contains inappropriate language, an error message is generated.

[1394] Step 7:

[1395] If an inappropriate phrase is detected during the automated review process, the server generates an error message and sends it to the terminal, where the user receives the error message and corrects it to an appropriate phrase.

[1396] Step 8:

[1397] Phrases that pass the screening are sent by the server to the emotion engine, which estimates the user's emotions from the images and audio data captured by the user's camera and microphone.

[1398] Step 9:

[1399] The emotion engine uses facial recognition and voice analysis technologies to estimate emotions from the user's facial expressions and tone of voice. For example, a smile is recognized as a positive emotion, while tears are recognized as a negative emotion.

[1400] Step 10:

[1401] The recognized emotion information is sent to a server and integrated with the speech generation means, which provides the emotion information to the speech generation means so that phrases are generated with a tone and inflection appropriate to the emotion.

[1402] Step 11:

[1403] The phrases and emotional information that pass the screening process are then fed into a generative AI model, which uses machine learning algorithms to generate audio data that mimics the voice of the designated talent or character, adding tone and inflection that corresponds to the recognized emotion.

[1404] Step 12:

[1405] The generated voice data is then sent back to the server and stored in a database, allowing users to access the voice data at any time later.

[1406] Step 13:

[1407] The server generates a link or file URL to send the audio data to the user's device, which then displays the link or URL in a user interface.

[1408] Step 14:

[1409] Users can click the link or file URL to play or download the generated audio data, allowing them to enjoy a customized message in the voice of the selected talent or character.

[1410] Step 15:

[1411] Users can share the generated voice message directly from their device to social media or messaging applications, for example, using LINE or Twitter.

[1412] Example 2

[1413] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1414] Conventional voice message generation systems have difficulty generating voice messages that reflect the user's emotions, preventing them from providing a personalized experience. Furthermore, the process for evaluating the appropriateness of phrases is insufficient, which can result in inappropriate phrases being generated. Furthermore, when imitating the voices of specific celebrities or characters, natural tones and intonations are often lacking.

[1415] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1416] In this invention, the server includes input means for a user to input a phrase, transmission means for transmitting the input phrase to the server, review means for automatically reviewing the transmitted phrase, voice generation means for converting a phrase that passes the review into a specified voice, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for recognizing emotions, and means for adjusting the tone and intonation of the voice in accordance with the emotion recognized by the emotion recognition means. This makes it possible to generate a personalized voice message that reflects the user's emotions and provide a more natural and appropriate voice message.

[1417] An "input means" is a device or software that provides an interface for a user to input a phrase.

[1418] "Transmission means" refers to a device or software that has the function of transmitting phrases and information entered by the user to the server.

[1419] "Screening means" refers to a device or software that has the function of automatically screening transmitted phrases and evaluating their appropriateness and the presence of prohibited words.

[1420] "Speech generation means" refers to a device or software that converts phrases that have passed the screening into a specified voice and generates voice data.

[1421] "Emotion recognition means" refers to devices or software that use algorithms to recognize emotions from the user's facial expressions and voice via the user's camera or microphone.

[1422] The "tone and intonation adjusting means" refers to a device or software that has the function of adjusting the tone and intonation of the generated voice in accordance with the user's emotion recognized by the emotion recognition means.

[1423] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. The system further incorporates an emotion engine that recognizes the user's emotions. Components of the system include a user, a terminal, a server, and the emotion engine.

[1424] A user inputs a desired phrase using their own device and sends it. A browser or a dedicated application is installed on the device, which functions as a user interface. The user interface includes a phrase input field, a talent or character selection menu, a send button, and a camera and microphone access permission screen.

[1425] The device generates a request containing the phrase entered by the user, information about the selected talent or character, and the emotion recognition results from the emotion engine, and sends this request to the server as an HTTP POST request.

[1426] The server receives this request and stores it in a database. The stored data contains information necessary for subsequent processing. The server also subjects the received phrases to an automated review process. This review process uses natural language processing (NLP) techniques to evaluate the appropriateness of the phrase. For example, the phrase "Happy Birthday" is deemed appropriate, but an error message is generated if it contains inappropriate language.

[1427] Phrases that pass the screening process are sent to the emotion engine for emotion recognition. The emotion engine uses the user's camera and microphone to apply algorithms that infer emotions from facial expressions and vocal tone. For example, if the user is smiling, it will be recognized as a positive emotion, and if they look sad, it will be recognized as a negative emotion.

[1428] The emotion information recognized by the emotion engine is sent to a server, which then connects to a voice generator. The server then inputs the selected phrases and emotion information into a generative AI model, which uses machine learning algorithms to not only imitate the voice of the designated talent or character, but also add tone and intonation according to the recognized emotion.

[1429] As an example of a specific prompt, let's say the user inputs the phrase "Happy birthday! Stay healthy!" in Talent A's voice while smiling. The emotion engine recognizes the user's smile as a positive emotion, and generates a bright-toned voice that reflects that emotion. Here's a concrete example of how this prompt can be input into the generative AI model:

[1430] "Generate the phrase 'Happy birthday! Stay healthy!' in Talent A's voice based on the positive emotions expressed by smiling users."

[1431] The generated audio data is returned to the server and stored in a database. This storage allows the user to access the audio data at any time later. The server generates a link or file URL to send the audio data to the user's device, and this link or URL is sent to the device. The user can play or download the audio data by clicking the link or URL displayed in the user interface.

[1432] Users can also share the generated voice messages directly from their devices on social media and messaging applications. For example, they can share messages with friends and family using LINE or Twitter. This allows users to easily and quickly enjoy customized messages in the voices of specific celebrities or characters. The built-in emotion engine provides more personalized voice messages that reflect the user's emotions.

[1433] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1434] Step 1: User enters phrase

[1435] The user accesses the system via the device's browser or a dedicated application. The user selects a phrase to enter in the phrase input field displayed in the user interface and a talent or character from a selection menu. Specifically, the user enters the phrase "Happy birthday! Stay healthy!" and selects Talent A. The entered data is stored internally on the device as a generated prompt.

[1436] input:

[1437] Phrase: "Happy birthday! Stay healthy!"

[1438] Talent or Character: Talent A

[1439] output:

[1440] Generated prompt data

[1441] Step 2: Generate and send a request

[1442] The device generates request data using the entered phrase and information about the selected talent or character, and sends it to the server as an HTTP POST request.

[1443] input:

[1444] Prompt Data

[1445] Data processing and calculation:

[1446] Generating an HTTP POST Request

[1447] output:

[1448] Sending requests to the server

[1449] Step 3: Data received and stored by the server

[1450] The server receives the request sent from the device and stores it in a database, including phrases and talent or character information.

[1451] input:

[1452] Send Request

[1453] Data processing and calculation:

[1454] Saving to a database

[1455] output:

[1456] Phrases and talent information stored in the database

[1457] Step 4: Automated Review Process

[1458] The server automatically reviews stored phrases for appropriateness using natural language processing (NLP) techniques: for example, "Happy Birthday" is deemed appropriate, while inappropriate phrases generate an error message.

[1459] input:

[1460] Phrases in the database

[1461] Data processing and calculation:

[1462] Appropriateness assessment using natural language processing

[1463] output:

[1464] Review result (appropriate or inappropriate)

[1465] Step 5: Emotion Recognition Process

[1466] Phrases deemed appropriate are sent to the emotion engine, which recognizes emotions from facial expressions and tone of voice via the user's camera or microphone. For example, if the user is smiling, it is recognized as a positive emotion.

[1467] input:

[1468] Phrase (reviewed)

[1469] User facial expression data

[1470] User's voice tone

[1471] Data processing and calculation:

[1472] Applying emotion recognition algorithms

[1473] output:

[1474] Recognized emotional information

[1475] Step 6: The speech generation process

[1476] Based on the emotional information from the emotion engine and the screened phrases, the server uses a generative AI model to generate audio in the voice of the designated talent or character, with tone and inflection adjusted according to the recognized emotion.

[1477] input:

[1478] phrase

[1479] Recognized emotional information

[1480] Data processing and calculation:

[1481] Speech generation using generative AI models

[1482] output:

[1483] Generated audio data

[1484] Step 7: Send and save the audio data

[1485] The generated voice data is returned to the server and stored in a database. This allows the user to access the voice data at any time later. The server generates a link or file URL to send the voice data to the user's device and sends this link to the user.

[1486] input:

[1487] Generated audio data

[1488] Data processing and calculation:

[1489] Saving to a database

[1490] Generate a download link or URL

[1491] output:

[1492] Links or URLs sent to users

[1493] Step 8: Sharing Function

[1494] Users receive the generated audio message and can click a link or URL to play or download the audio data, or share the audio message with friends and family via social media or messaging applications.

[1495] input:

[1496] Links or URLs sent to users

[1497] Data processing and calculation:

[1498] Download or play audio data

[1499] Share audio messages

[1500] output:

[1501] Playing audio data

[1502] Share audio messages

[1503] (Application example 2)

[1504] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1505] In modern advertising, sending unilateral messages without considering user emotions cannot be expected to be effective. Furthermore, uniform voice messages have difficulty attracting user attention, and a lack of personalization leads to reduced engagement with advertising. Therefore, there is a need for a system that can grasp a user's emotional state in real time and generate and deliver personalized advertising messages with appropriate voice tones.

[1506] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for the user to input a phrase, a transmission means for transmitting the input phrase to the server, an examination means for automatically examining the transmitted phrase, an emotion recognition means for recognizing an emotion, a voice generation means for converting the phrase into a specified voice with an audio tone corresponding to the emotion recognized by the emotion recognition means, and a transmission means for transmitting the generated voice data to the user's terminal and playing it back. This makes it possible to generate and deliver a personalized voice message that takes the user's emotions into consideration.

[1507] "User" means a person or entity that utilizes the system to input phrases and generate voice messages.

[1508] "Input means" refers to a device or software function that allows a user to input a desired phrase.

[1509] "Transmission means" refers to a device or software function for transmitting the input phrase to the server.

[1510] A "screening tool" is a process or algorithm that automatically determines the appropriateness of a submitted phrase.

[1511] "Emotion recognition means" refers to a device or software function that reads and analyzes a user's emotions through a camera or microphone.

[1512] A "voice generation means" is a device or software function that converts phrases into voice data based on a specified voice or emotional tone.

[1513] The "server" is a central processing unit that processes various data and transmits audio data and control information to client terminals.

[1514] A "terminal" is a device (such as a smartphone or PC) that a user directly operates and interacts with the system.

[1515] A "phrase" is a series of words or sentences that a user wishes to convert into a voice message.

[1516] This invention relates to a system that allows users to create and use custom voice messages using the voice of a specific celebrity or character. This system incorporates an emotion engine that recognizes the user's emotions, enabling the creation of more personalized voice messages.

[1517] composition

[1518] This system consists of a user terminal, a server, an emotion recognition engine, and a voice generation engine. The user terminal is assumed to be a smartphone, but other devices can also be used. The server acts as a central processing unit and manages the emotion recognition and voice generation processes.

[1519] Hardware / Software used

[1520] Hardware:

[1521] Smartphone (iOS, Android)

[1522] software:

[1523] Microsoft Azure's Face API or Google Cloud's Vision AI (emotion recognition)

[1524] OpenAI's GPT-3 (Speech Generation)

[1525] HTTP server (data transmission and reception)

[1526] Operation method and configuration

[1527] 1. User Interface:

[1528] The user launches the app and is asked for camera and microphone permissions, which initiates the emotion recognition process. The user then enters the desired message in the phrase input field and presses send.

[1529] 2. Emotion recognition:

[1530] Data is collected in real time from the camera and microphone on the user's device and sent to the emotion recognition API, where the user's facial expressions and tone of voice are analyzed to obtain emotional information.

[1531] 3. Data transmission:

[1532] The acquired emotion information and the input phrase are sent to the server using an HTTP POST request.

[1533] 4. Automated Review:

[1534] The server receives the submitted phrase and evaluates its appropriateness using natural language processing (NLP) techniques. If it contains inappropriate language, it returns an error message; if it is an appropriate phrase, it proceeds to the next step.

[1535] 5. Speech generation:

[1536] Based on the phrases and emotional information that pass the screening process, a speech generation engine (OpenAI's GPT-3) is used to generate a custom voice message in the voice of the designated talent or character, with the tone and intonation adjusted according to the recognized emotion.

[1537] 6. Data transmission and playback:

[1538] The generated audio data is then sent back to the user's device via the server, where it can be played within the application, and a sharing link is also generated, allowing the user to share the audio message on social media or via messaging apps.

[1539] Specific examples

[1540] For example, if a user is in a good mood and types the phrase "Thank you for your hard work today!", the emotion recognition engine will recognize the positive emotion and generate a message in a bright tone such as "Thank you for your hard work today!"

[1541] Example prompt sentence:

[1542] "Please generate an advertising message in the voice of celebrity X to be played when the user has a happy expression. The message should introduce product Z in a cheerful tone."

[1543] In this way, it becomes possible to generate and deliver personalized voice messages that take into account the user's emotions.

[1544] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1545] Step 1:

[1546] The user launches the application and grants permission to use the camera and microphone. Once this is complete, the system is ready to collect facial and voice data in real time.

[1547] Input: App launch, camera and microphone permissions

[1548] Output: Real-time facial expression and voice data acquisition status

[1549] Step 2:

[1550] The user enters the phrase using an input device and presses the send button, at which point the phrase is saved on the device and ready to be sent in the next step.

[1551] Input: A phrase entered by the user

[1552] Output: The phrase you entered is saved to your device.

[1553] Step 3:

[1554] The user device sends the phrase entered and the facial expression and voice data acquired in real time to the server as an HTTP POST request. The transmitted data includes the phrase entered by the user, facial expression data, and voice data.

[1555] Input: User-entered phrases, facial expressions and voice data captured in real time

[1556] Output: Request data sent to the server

[1557] Step 4:

[1558] The server then puts the received phrases through an automated review process. The review uses natural language processing (NLP) technology to evaluate the appropriateness of the phrase. If the phrase contains inappropriate content, an error message is generated and returned to the terminal. If the phrase is appropriate, the process proceeds to the next step.

[1559] Input: Phrase sent to the server

[1560] Output: Evaluation result for suitability, error message or approval flag

[1561] Step 5:

[1562] The server sends the facial and voice data acquired in real time to the emotion recognition device, which then analyzes the data and recognizes the user's emotions (such as Microsoft Azure's Face API or Google Cloud's Vision AI). The recognized emotion data is then sent to the server.

[1563] Input: Real-time facial and voice data

[1564] Output: Recognized emotion data

[1565] Step 6:

[1566] Based on the phrases and recognized emotion data that have passed the screening process, the server uses a voice generation engine (OpenAI's GPT-3) to generate a custom voice message in the voice of the specified talent or character, taking into account emotional tone during the generation process.

[1567] Input: phrases that passed the screening, recognized emotion data

[1568] Output: Generated audio data

[1569] Step 7:

[1570] The server sends the generated audio data to the user terminal. In this step, the generated audio data and a playback link are sent to the user terminal.

[1571] Input: Generated audio data

[1572] Output: Audio data and playback link sent to the user's device

[1573] Step 8:

[1574] The user's device will then play the received audio data, allowing the user to listen to it. Additionally, the user will have the ability to share the audio message on social media or messaging apps.

[1575] Input: Audio data and playback link sent to the user's device

[1576] Output: Played audio data and a shareable link

[1577] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1578] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1579] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1580] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1581] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1582] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1583] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1584] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1585] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1586] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1587] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1588] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1589] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1590] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1591] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1592] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1593] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1594] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1595] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1596] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1597] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1598] The following is further disclosed regarding the above embodiment.

[1599] (Claim 1)

[1600] an input means for a user to input a phrase;

[1601] a transmitting means for transmitting the input phrase to a server;

[1602] a reviewing means for automatically reviewing the submitted phrases;

[1603] a voice generating means for converting the phrases that have passed the screening into a designated voice;

[1604] a transmitting means for transmitting the generated voice data to a user terminal;

[1605] A system including:

[1606] (Claim 2)

[1607] 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

[1608] (Claim 3)

[1609] 2. The system of claim 1, wherein the voice generation means uses a machine learning algorithm to imitate the voice of a specific target.

[1610] (Claim 4)

[1611] 10. The system of claim 1, further comprising means for storing the generated speech data in a database and providing it in a user-accessible format.

[1612] (Claim 5)

[1613] 10. The system of claim 1, further comprising a sharing means for enabling a user to share the generated voice message on social media or messaging applications.

[1614] "Example 1"

[1615] (Claim 1)

[1616] an input means for a user to input a phrase;

[1617] a transmitting means for transmitting the input phrase to a server;

[1618] a reviewing means for automatically reviewing the submitted phrases;

[1619] a voice generating means for converting the phrases that have passed the screening into designated voices;

[1620] a storage means for storing the generated voice data in a database and generating a link or file URL for the voice data;

[1621] a sending means for sending the generated link or file URL to the user's terminal;

[1622] playback means for a user to play or download the generated audio data;

[1623] A system including:

[1624] (Claim 2)

[1625] 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

[1626] (Claim 3)

[1627] 2. The system of claim 1, wherein the voice generation means uses a machine learning algorithm to mimic the voice of a specific target.

[1628] "Application Example 1"

[1629] (Claim 1)

[1630] an input means for a user to input a phrase;

[1631] a transmitting means for transmitting the input phrase to a server;

[1632] a reviewing means for automatically reviewing the submitted phrases;

[1633] a voice generating means for converting the phrases that have passed the screening into a designated voice;

[1634] a transmitting means for transmitting the generated voice data to a user terminal;

[1635] a message generating means for generating a specific status notification message using the generated voice data;

[1636] playback means for allowing a user to play or share a particular status notification message;

[1637] A system including:

[1638] (Claim 2)

[1639] 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

[1640] (Claim 3)

[1641] 2. The system of claim 1, wherein the voice generation means uses a machine learning algorithm to imitate the voice of a specific target.

[1642] "Example 2: Combining Emotion Engines"

[1643] (Claim 1)

[1644] an input means for a user to input a phrase;

[1645] a transmitting means for transmitting the input phrase to a server;

[1646] a reviewing means for automatically reviewing the submitted phrases;

[1647] a voice generating means for converting the phrases that have passed the screening into a designated voice;

[1648] a transmitting means for transmitting the generated voice data to a user terminal;

[1649] emotion recognition means for recognizing emotions;

[1650] means for adjusting the tone and intonation of the voice in response to the emotion recognized by the emotion recognition means;

[1651] A system including:

[1652] (Claim 2)

[1653] 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

[1654] (Claim 3)

[1655] 2. The system of claim 1, wherein the voice generation means uses a machine learning algorithm to mimic the voice of a specific target and adjusts the tone and intonation of the voice according to the emotion recognized by the emotion recognition means.

[1656] "Application example 2 when combining emotion engines"

[1657] (Claim 1)

[1658] an input means for a user to input a phrase;

[1659] a transmitting means for transmitting the input phrase to a server;

[1660] a reviewing means for automatically reviewing the submitted phrases;

[1661] emotion recognition means for recognizing emotions;

[1662] a voice generating means for converting a phrase into a designated voice with a voice tone corresponding to the emotion recognized by the emotion recognition means;

[1663] a transmitting means for transmitting the generated voice data to a user's terminal and playing the voice data;

[1664] A system including:

[1665] (Claim 2)

[1666] 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

[1667] (Claim 3)

[1668] 2. The system of claim 1, wherein the voice generation means uses a machine learning algorithm to imitate the voice of a specific target. [Explanation of symbols]

[1669] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. an input means for a user to input a phrase; a transmitting means for transmitting the input phrase to a server; a reviewing means for automatically reviewing the submitted phrases; a voice generating means for converting the phrases that have passed the screening into a designated voice; a transmitting means for transmitting the generated voice data to a user terminal; A system including:

2. 2. The system of claim 1, wherein the reviewing means uses natural language processing to evaluate the appropriateness of the phrases.

3. 10. The system of claim 1, wherein the voice generating means uses a machine learning algorithm to mimic the voice of a specific target.

4. 10. The system of claim 1, further comprising means for storing the generated speech data in a database and providing it in a user accessible format.

5. 10. The system of claim 1, further comprising a sharing means for enabling a user to share the generated voice message on social media or messaging applications.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A