system
The system addresses the inefficiencies in manual intonation adjustment by using a deep learning model with a feedback loop to generate high-quality, natural-sounding speech in voice dialogue systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-10
- Publication Date
- 2026-06-22
AI Technical Summary
Existing voice dialogue systems require manual adjustment for natural and emotional intonation, incurring significant time and cost, and lack efficient mechanisms for reflecting user feedback to improve voice quality.
A system that acquires speech feature data, preprocesses it, and uses a deep learning model to generate speech with natural intonation, incorporating a feedback loop for model improvement.
Enables efficient and high-quality speech generation with natural intonation by continuously refining the model based on user feedback, enhancing user satisfaction.
Smart Images

Figure 2026101165000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0005] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the current voice dialogue system, in order to impart a natural and emotional intonation, manual adjustment is required, and there is a problem that a huge amount of time and cost are incurred in the process. In addition, there is a lack of a system for efficiently reflecting the variation in the quality of the generated voice and the user feedback. This invention aims to provide a method for efficiently performing natural and reliable voice generation for these problems that are difficult to solve by existing methods.
Means for Solving the Problems
[0005] This invention provides a means for acquiring speech feature data from a database and preprocessing it. Furthermore, it proposes a system that includes a means for learning the optimal intonation pattern for text using a deep learning model. The system has means for generating speech data with natural intonation using the trained model based on text input by the user, and it is possible to play the generated speech back to the user and receive their evaluation. In addition, by constructing a feedback loop to modify the model based on the evaluation and improve speech quality, efficient and high-quality speech generation is achieved.
[0006] A "database" is an information system built to store data and allow for efficient retrieval and management.
[0007] "Speech feature data" refers to data that represents acoustic characteristics extracted from speech waveforms, and plays an important role in speech recognition and generation.
[0008] "Means of preprocessing" refer to processing methods or devices for converting raw data into a format suitable for analysis or learning.
[0009] A "deep learning model" is an algorithm based on a multi-layered neural network, designed to perform complex pattern recognition and prediction.
[0010] "Intonation patterns" refer to changes in pitch, volume, and speed of speech, and are phonetic attributes used to convey emotions and meaning.
[0011] "Adding natural intonation" is the process of adding a natural vocal tone to the generated speech that closely resembles human speech.
[0012] "Means for generating audio data" refers to devices or methods for creating audio signals from text information.
[0013] A "feedback loop" is a cyclical information processing mechanism that uses the output results to improve a system or process. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention provides a system for automating the application of natural intonation in a voice dialogue system. This system primarily consists of three elements: a server, a terminal, and a user.
[0036] First, the server retrieves historical audio feature data from the company's internal database and preprocesses it into a format suitable for machine learning. The audio data is converted into acoustic features such as MFCCs and spectrograms. This conversion extracts important information from the audio waveform, which can then be used to train the AI model.
[0037] Next, the server builds a deep learning model for the dialogue system and trains it using pre-processed data. This model employs a multi-layer neural network architecture and has the ability to generate speech with natural intonation based on the input text. During the training phase, techniques such as cross-validation are used to improve the accuracy of intonation and expression patterns.
[0038] The user inputs arbitrary text through the terminal. This text data is sent from the terminal to the server, which uses a trained model to generate speech data with natural intonation corresponding to the meaning and context of the text.
[0039] The generated audio data is sent back to the terminal and played back by the user. During this process, the user evaluates whether the generated audio sounds natural and provides feedback as needed. This feedback information is sent to the server and used to further improve the model's accuracy. This feedback loop allows the system to continuously improve, resulting in higher quality generated audio.
[0040] For example, if a user types "Tell me the weather for tomorrow," the server can generate a friendly voice response with an emotional intonation, such as "It's supposed to be sunny tomorrow." This allows the user to have a natural, near-human communication experience.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The server retrieves previously collected audio data from the database. This data is stored as pairs of text and corresponding audio files.
[0044] Step 2:
[0045] The server performs preprocessing on the audio data. Specifically, it extracts acoustic features from the audio waveform and converts them into formats such as MFCCs and spectrograms. This makes it easier to represent the characteristics of the audio digitally.
[0046] Step 3:
[0047] The server uses the extracted features to train a deep learning model. This model uses manually adjusted speech data as training data to learn the optimal intonation patterns for text.
[0048] Step 4:
[0049] The user uses the terminal to input text they want to interact with. For example, they might type an everyday question.
[0050] Step 5:
[0051] The terminal sends the entered text to the server. The transmitted data contains the text information necessary for speech generation.
[0052] Step 6:
[0053] The server generates speech data with natural intonation based on the received text, using a trained AI model. During this process, it adjusts pitch, rhythm, and emphasis according to the context of the text.
[0054] Step 7:
[0055] The server sends the generated audio data to the terminal. The data sent includes an audio file in a format playable by the user.
[0056] Step 8:
[0057] The device plays the generated audio data to the user. The user can then check the naturalness and intonation of the audio.
[0058] Step 9:
[0059] Users provide feedback on the played audio, evaluating whether it is appropriate and whether it needs correction.
[0060] Step 10:
[0061] The device sends user feedback to the server. This feedback is used to adjust and improve the model.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] In voice dialogue systems, there is a need to generate speech with natural and user-friendly intonation. However, conventional technologies have suffered from insufficient intonation, resulting in low user satisfaction. Furthermore, there is a lack of mechanisms to adequately evaluate the naturalness of the generated speech and incorporate feedback to improve its quality.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for acquiring and converting acoustic feature data from a data storage device, means for learning appropriate intonation patterns for linguistic information using a multilayer learning model, and means for generating acoustic data with natural intonation using the trained model based on linguistic information input by the user. This provides users with natural and familiar speech and enables continuous model improvement through feedback.
[0067] A "data storage device" is a device that stores acoustic characteristic data and allows the data to be retrieved as needed.
[0068] "Acoustic feature data" refers to data obtained by analyzing audio information and representing it numerically in forms such as Mel-frequency cepstrum coefficients and spectrograms.
[0069] "Means of conversion" refers to a process or device that converts acquired acoustic data into a format that is easily processed by a machine learning model.
[0070] A "multilayer learning model" is a neural network with multiple layers, designed to learn various intonation patterns for linguistic information.
[0071] "Linguistic information" refers to data about text spoken as audio and its context.
[0072] An "appropriate intonation pattern" is a pattern that shows appropriately adjusted pitch, volume, and rhythm to achieve natural and friendly conversation.
[0073] A "pre-trained model" is a deep learning model that has been trained in advance using data and is tuned to handle a specific task.
[0074] "Natural intonation" refers to intonation that is similar to the sounds humans actually make in real conversations, and is intended to provide a natural conversational experience.
[0075] "Audio data" refers to digital audio information that can be output as sound.
[0076] "Evaluation information" refers to data that shows the content of feedback and evaluations that users provide to the generated audio.
[0077] "Means of tuning a model" refers to the process of making modifications or updates to a machine learning model based on evaluation information in order to improve its performance.
[0078] This invention includes a process in which a server, a terminal, and a user cooperate to generate speech with natural intonation, as part of a voice dialogue system.
[0079] The server accesses a data storage device that stores acoustic feature data and retrieves the necessary data. This data is converted into formats such as MFCC (Mel-Frequency Cepstral Coefficients) and spectrograms, which serve to extract important features from the speech. The conversion is performed using dedicated speech processing software, and the data is preprocessed into a format suitable for machine learning. LibROSA, a Python®-based library, is often used for this purpose.
[0080] Subsequently, the server uses deep learning frameworks such as TENSORFLOW® and PyTorch to build a multi-layer learning model. This model is trained using pre-collected acoustic data with the goal of learning appropriate intonation patterns. Natural language processing techniques are also used in conjunction with intonation learning to take into account the meaning and sentiment of the text.
[0081] The user inputs text into the system via a terminal device. The input text is processed by the terminal and sent to the server as text data. An example of text is the prompt "Tell me the weather tomorrow." This text is converted into speech data with natural intonation by a trained deep learning model. As a result, the user can receive a natural and immersive response that closely resembles human dialogue.
[0082] The generated audio data is sent back to the device and played back to the user through speakers or headphones. This playback allows the user to evaluate the intonation and naturalness of the generated audio and provide feedback. This feedback is entered through the user interface and sent to the server. This information is used to adjust and improve the model, contributing to improved quality in subsequent audio generation processes.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The server retrieves acoustic feature data from the data storage device. In this step, audio recording data is retrieved from the database and output as an audio file format (e.g., WAV).
[0086] Step 2:
[0087] The server preprocesses the acquired acoustic data. This process uses an audio processing library such as LibROSA to convert the audio recording data into MFCCs or spectrograms. As a result, it outputs acoustic features as numerical data usable in machine learning models.
[0088] Step 3:
[0089] The server trains a multi-layer learning model using pre-processed data. At this stage, training is performed using TensorFlow or PyTorch, and the output model learns appropriate intonation patterns that take into account pitch and rhythm.
[0090] Step 4:
[0091] The user enters text into the system via the terminal. The entered text is presented as an example prompt, "Tell me tomorrow's weather," and sent to the server as text data.
[0092] Step 5:
[0093] The server inputs the received text data into a trained model and generates audio data with natural intonation. Specifically, it analyzes the text and outputs audio files (e.g., WAV) that reflect the emotion and context.
[0094] Step 6:
[0095] The device receives the generated audio data and plays it for the user via its playback function. In this step, the audio is output and played back through the speaker or headphones.
[0096] Step 7:
[0097] Users evaluate the naturalness and intonation of the played audio and provide feedback. This feedback is entered as an evaluation statement and sent to the server as data for improvement.
[0098] Step 8:
[0099] The server adjusts the model based on the feedback it receives. This feedback is used in the next speech generation process, where data calculations are performed to improve the model's accuracy and produce a more effective model.
[0100] (Application Example 1)
[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0102] Conventional voice dialogue systems suffer from problems such as mechanical and unnatural intonation, making effective communication with users difficult. Furthermore, while there is a demand for natural voice dialogue in home electronic devices, current technology has its limitations. Therefore, there is a need for a means to achieve more natural and human-like voice dialogue.
[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0104] In this invention, the server includes means for acquiring voice feature information from an information storage device and preprocessing this information, means for learning the optimal speech pattern for a sentence using a deep learning model, and means for generating voice information with natural speech based on the sentence input by the user using the trained model. This makes it possible for home electronic devices to provide more natural and effective conversational functions.
[0105] An "information storage device" is a device used to store data such as audio and text, and is used for data retrieval and preprocessing as needed.
[0106] "Voice feature information" refers to the basic attributes and patterns extracted from speech data, and in particular, includes features that contribute to natural intonation.
[0107] "Preprocessing" refers to the process of converting acquired data into a format suitable for analysis and learning, and includes data normalization and feature extraction.
[0108] A "deep learning model" is a model that uses a neural network consisting of multiple layers to learn complex patterns in data.
[0109] The "optimal vocalization pattern" refers to the tone and intonation of the voice that is suitable for achieving a natural and human-like intonation in response to the input text.
[0110] A "user" refers to a person who interacts with this system and utilizes its functions using voice input or text input.
[0111] "Text" refers to the text data to be processed, including the content input for speech generation.
[0112] A "trained model" refers to a model that has been trained using deep learning to specialize in specific patterns or tasks.
[0113] "Voice information with added pronunciation" refers to audio data generated based on text, using natural intonation.
[0114] A "household machine" is an automated machine used within the home, designed to provide a variety of functions, including conversational capabilities.
[0115] "Dialogue functionality" refers to a function that enables humans and machines to exchange information and communicate through voice and text.
[0116] This invention relates to a system for generating natural intonation of speech data. This system is based on three main components: a server, a terminal, and a user.
[0117] The server retrieves voice feature information from its data storage device. This includes previously collected audio data, which is preprocessed to make it suitable for training deep learning models. Specifically, it converts audio waveforms into acoustic features such as MFCCs and spectrograms. This conversion process utilizes servers equipped with NVIDIA GPU devices and TensorFlow software.
[0118] Next, the server constructs a deep learning model, such as a recurrent neural network (RNN) or its variant, the LSTM (Long Short-Term Memory) model, and trains it using pre-processed data. The trained model has the ability to add natural intonation to text and generates speech based on the text data entered by the user.
[0119] Users access the system through a device that allows for voice and text-based interaction. This device can be a smartphone, tablet, or even a robot as a household appliance. Text data provided by the user is sent to the server, and the natural-sounding voice generated by the server is then provided back to the user through the device.
[0120] A concrete example would be a scenario where a household appliance gently informs the user, "Today's weather is sunny," while they are getting ready in the morning, and then prompts them to confirm their daily schedule. An example of a prompt corresponding to this specific example would be, "Generate a response to the user's question with natural intonation."
[0121] This allows users to experience more human-like communication through home electronic devices.
[0122] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0123] Step 1:
[0124] The server retrieves voice feature information from the information storage device. This information includes past audio data and is in a format that allows mapping to text. The data is obtained from the original audio file and used as input for the next preprocessing step.
[0125] Step 2:
[0126] The server preprocesses the acquired audio data. This preprocessing includes conversion to MFCCs and spectrograms. This allows for the extraction of temporal and frequency features of the audio, resulting in high-dimensional audio features. The output is provided as a vector of acoustic features.
[0127] Step 3:
[0128] The server uses pre-processed speech features as input to build and train a deep learning model. Specifically, it uses a recurrent neural network (RNN) or LSTM model to learn the relationship between speech and text. In this step, the model's weights are updated to acquire the optimal speech pattern. Once training is complete, the model is ready to generate speech.
[0129] Step 4:
[0130] The user inputs text through their device and sends that data to the server. The input text data is provided as a sentence to which intonation should be added.
[0131] Step 5:
[0132] The server receives text sent by the user and generates speech using a trained model. The generation process converts the input text into speech data with optimal intonation. The output speech is then sent to the terminal as a digital audio file.
[0133] Step 6:
[0134] The device receives audio data sent from the server and plays it back to the user. This allows the user to hear audio with natural intonation.
[0135] Step 7:
[0136] Users provide feedback on the generated audio. This feedback concerns the naturalness and intelligibility of the audio and is sent to the server and recorded in a database.
[0137] Step 8:
[0138] The server readjusts the model based on user feedback. It analyzes the feedback and uses cross-validation to improve the model's accuracy. This step improves the overall system quality.
[0139] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0140] This invention provides a system for achieving natural intonation and user-centric dialogue in a voice dialogue system. This system consists of four elements: a server, a terminal, a user, and an emotion engine.
[0141] First, the server retrieves speech feature data from the database and preprocesses it. The speech data is processed into acoustic features and converted into a format suitable for training with a deep learning model. Intonation patterns are then learned based on this.
[0142] Next, during the training process using a deep learning model, the server learns the optimal intonation pattern for the user's input text. An emotion engine is incorporated, which can recognize emotions from the user's voice input and text. This engine uses emotion analysis technology to analyze keywords in the text and speech parameters to identify the user's emotional state.
[0143] When a user interacts using a device, the device sends data to the server in response to the user's input (text or voice). The data received by the server is then generated as speech data with natural intonation using a trained model. This generation process also takes into account the user's emotional state, and speech adjustments are made to suit that emotion.
[0144] The generated audio is sent to the device and played back to the user. At this time, the intonation adjusted by the emotion engine provides the user with a more natural and personalized conversational experience. Through this audio playback, the emotional nuances of the speech are more easily conveyed to the user.
[0145] Furthermore, user feedback and evaluations of the voice are sent from the device to the server. This feedback is shared with the emotion engine and used to improve the accuracy of the model and the engine. For example, if a user inputs "I'm feeling a little tired today," the emotion engine recognizes "fatigue," and the server generates a response with an intonation that takes this into consideration. This allows the user to experience more friendly and comfortable communication.
[0146] Thus, the present invention is a system that uses an emotion engine to realize natural voice dialogue optimized for the user.
[0147] The following describes the processing flow.
[0148] Step 1:
[0149] The server retrieves past audio data and its associated emotion labels from the database. This data is stored as text, corresponding audio files, and emotion information.
[0150] Step 2:
[0151] The server extracts acoustic features from the audio data and trains a deep learning model with emotion labels. The model learns the relationship between intonation patterns and emotional states in relation to speech, creating a foundation for natural speech generation.
[0152] Step 3:
[0153] The user inputs the text they want to speak through the device. Sometimes, the user's voice is also input at the same time.
[0154] Step 4:
[0155] The terminal sends user input text and voice to the server. An emotion engine also operates at this point to analyze the emotion in the voice.
[0156] Step 5:
[0157] The server uses an emotion engine to recognize emotions from the user's voice and adjusts the intonation generated by the deep learning model based on that emotion information.
[0158] Step 6:
[0159] The server generates audio data with the optimal intonation for the user's text. The generated audio takes into account the user's emotional state.
[0160] Step 7:
[0161] The server sends the generated voice data to the terminal. The voice, adjusted with emotional information, is more natural and approachable.
[0162] Step 8:
[0163] The device plays the generated audio to the user, who then receives feedback from the system.
[0164] Step 9:
[0165] Users provide feedback on the played audio, evaluating whether it is appropriate and what needs improvement.
[0166] Step 10:
[0167] The device sends user feedback to the server. This feedback is used for the continuous improvement of the emotion engine and deep learning models.
[0168] (Example 2)
[0169] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0170] In voice dialogue systems, generating voice responses with natural intonation that takes into account the user's emotions has been difficult with conventional technologies. Furthermore, there is a need for a system that effectively utilizes user feedback to improve model accuracy.
[0171] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0172] In this invention, the server includes means for preprocessing voice information obtained from a database, means for learning the optimal intonation pattern for the information using a machine learning algorithm, means for generating voice information with natural intonation using a trained model based on information input by the user, means for identifying the emotional state from the input information using sentiment analysis technology, and means for modifying the model based on the evaluation and emotional state identification results. This enables the generation of natural voice responses that correspond to the user's emotional state and continuous model improvement.
[0173] A "database" is a storage or management system that systematically organizes information to facilitate searching and access.
[0174] "Vocal information" refers to data related to human speech, and includes acoustic features, intonation, and parameters indicating emotion.
[0175] "Preprocessing" is the stage of shaping or transforming raw data into a format suitable for subsequent processing, and involves denoising data and extracting features.
[0176] A "machine learning algorithm" is a set of logic and procedures that allows a computer to learn patterns from data and perform predictions and classifications.
[0177] An "intonation pattern" is a structure that indicates the placement of intonation and accent in speech, and is important for achieving natural speech.
[0178] A "trained model" is an algorithm that has undergone an initial training process and incorporates knowledge based on diverse data.
[0179] "Sentiment analysis technology" is a technology that automatically identifies and classifies a person's emotions expressed in text or audio.
[0180] "Evaluation" refers to the act of users reviewing the results generated by a system and providing feedback based on quality and satisfaction.
[0181] "Modification" refers to making adjustments or changes to a system or algorithm in order to improve its initial state or its state after changes.
[0182] As an embodiment of this invention, details of a dialogue system using voice information are provided. The system consists of a server, terminals, users, and a data communication network connecting them.
[0183] The server acts as a central processing unit, handling audio information processing. Specifically, it retrieves audio information from the database and performs preprocessing. This preprocessing involves using audio processing libraries (e.g., Librosa) to remove noise and extract acoustic features. These features are then processed by deep learning models (e.g., RNNs and Transformers) to learn intonation patterns. The server also incorporates sentiment analysis technology, which uses natural language processing tools (e.g., NLTK and TextBlob) to identify emotions.
[0184] The user inputs voice or text via a terminal acting as an interactive device. The user's input is transmitted to the server through the terminal. The terminal then forwards the received voice data or text information directly to the server.
[0185] The device receives voice information generated from the server. This voice information is generated with intonation that takes the user's emotions into consideration, providing a more natural conversational experience. The user can provide feedback on this played voice information through the device.
[0186] The server accumulates user evaluations and feedback, and continuously improves the generated AI model. Cross-validation technology can be used to improve the model's accuracy.
[0187] As a concrete example, here is an example of a prompt: "Based on the text received from the user, generate a voice response with natural intonation that takes into account the emotional state." This prompt enables the system to engage in highly engaging conversations with the user, providing efficient and satisfying communication.
[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0189] Step 1:
[0190] The server retrieves audio information from the database. It receives audio samples and associated metadata as input. This audio information contains various intonation patterns and is the target of analysis. The server uses an audio processing library to extract fundamental frequencies and spectral information and convert them into acoustic features. The output provides acoustic features for learning intonation.
[0191] Step 2:
[0192] The server inputs the acquired acoustic features into a deep learning model to learn the optimal intonation pattern. Machine learning algorithms are used here. During this process, the model acquires the ability to recognize patterns in information and identify intonation patterns. The learned intonation pattern is provided as output.
[0193] Step 3:
[0194] The user inputs data via voice or text through the device. This input includes data that reflects the user's intentions and emotions. The device sends this data to the server in preparation for the next processing step.
[0195] Step 4:
[0196] The server uses sentiment analysis technology to identify the user's emotional state based on input data received from the user. Input can be text or audio, along with a currently learned model. Here, natural language processing tools are used to analyze keywords and phrases to determine the user's emotions. The identified emotional state is provided as output.
[0197] Step 5:
[0198] The server uses a trained model to adjust intonation according to the identified emotional state and generate a voice response. The input consists of trained intonation patterns and the user's emotion. The output is voice data with a natural intonation that matches the user's emotion.
[0199] Step 6:
[0200] The device receives the generated audio and plays it for the user. It handles audio data sent from the server as input. During playback, it ensures that the intonation is natural and aligns with the user's emotions. The final audio experience delivered to the user is then produced as output.
[0201] Step 7:
[0202] Users provide feedback on voice responses. As input, they enter their evaluation or comments on the generated voice into the device. The device sends this to the server, which helps improve the model. As output, the feedback is accumulated and used to generate the next response.
[0203] (Application Example 2)
[0204] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0205] Conventional voice dialogue systems have struggled to provide natural intonation that fully considers the user's emotions, making it difficult to achieve emotionally resonant dialogue. This problem makes it challenging to provide friendly and comfortable communication for users.
[0206] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0207] In this invention, the server includes means for acquiring speech feature information from a data set and preprocessing it, means for learning the optimal intonation pattern for the recording using a machine learning system, and means for analyzing the user's emotions and reflecting emotion-based intonation in the speech information. This makes it possible to provide emotion-adaptive speech dialogue with natural intonation.
[0208] A "data set" is a collection of data that is gathered and managed for a specific purpose.
[0209] "Audio feature information" refers to information containing features such as frequency, intensity, and pitch extracted from audio data.
[0210] "Preprocessing" is the process of transforming and processing data to make it suitable for analysis and learning.
[0211] A "machine learning system" is a system that uses algorithms to learn rules and patterns from data and perform predictions and classifications.
[0212] "Records" are data or documents that are stored and protected for later reference.
[0213] An "intonation pattern" is a pattern that shows the combination and arrangement of intonation in the sounds of a language.
[0214] A "user" is a person who uses a system or device.
[0215] "Emotional analysis" is the process of extracting emotional states from audio or text as numerical values or labels.
[0216] "Audio information" refers to all information transmitted as sound, including audio signal data.
[0217] "Emotionally adaptive voice interaction" refers to voice interaction that recognizes the user's emotions and engages in natural conversation tailored to those emotions.
[0218] The system implementing this invention consists of three elements: a server, a terminal, and a user. The server acquires speech feature information from a data set and preprocesses that data. The preprocessed data is used by a machine learning system to learn the optimal intonation pattern for recording. In particular, TensorFlow / Keras is used as a deep learning framework to learn natural intonation. Speech processing tools such as PyDub are utilized for processing the speech data.
[0219] The server then generates speech information with natural intonation using a trained system based on the user's input. The terminal receives the user's input and sends that information to the server. The server analyzes the user's emotions using an emotion recognition engine. Using the natural language processing library NLTK, it determines the user's emotions and generates speech information with the appropriate intonation.
[0220] The generated voice information is sent to the terminal and played back by the user. Through the played voice, emotionally adaptive voice dialogue is realized. This dialogue can be used particularly in home assist robots and improves the quality of everyday communication.
[0221] As a specific example, in a morning scenario, when the robot says, "Good morning. How are you feeling today?", and the user inputs, "I'm feeling a little down," the robot will provide an emotionally adaptive response. In this case, the robot will reply in a gentle tone of voice, "I see. Let's do something to help you relax."
[0222] An example of a prompt is, "Generate a kind and considerate intonation pattern for when a user says, 'Today was tough.'" This sentence is used in a generative AI model to generate responses appropriate to the user's emotions.
[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0224] Step 1:
[0225] The terminal acquires text and voice data entered by the user. This data is sent to the server as input data. The input data is expected to contain the user's intentions and emotions.
[0226] Step 2:
[0227] The server extracts speech feature information from the received speech data and performs natural language processing on the text data as needed. It extracts acoustic features from the speech data and performs emotion recognition on the text data. Based on the input data, it analyzes the features of the speech data and generates a feature vector as an intermediate output.
[0228] Step 3:
[0229] The server uses a deep learning model to generate the optimal intonation pattern from the feature vector. Here, the learned pattern is applied using a generative AI model. The output of this step is the generated intonation pattern.
[0230] Step 4:
[0231] The server uses an emotion recognition engine to analyze the user's emotions and adjust the intonation accordingly. Specifically, it takes the emotion analysis results as input and adjusts the intonation appropriately. It generates speech data with intonation that matches the emotion.
[0232] Step 5:
[0233] The server sends the generated audio data to the terminal. The terminal plays this data back to the user. This provides the user with natural and emotionally appropriate voice responses. The final output is the intoned audio played back to the user.
[0234] Step 6:
[0235] The user evaluates the played audio and inputs their feedback into the device. The device then sends the feedback to the server. This evaluation information is used to further improve the system.
[0236] Step 7:
[0237] Based on the feedback received, the server adjusts the speech generation system and improves the accuracy of the learning model and emotion engine in preparation for the next use. This improves the overall performance of the system.
[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0241] [Second Embodiment]
[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0254] This invention provides a system for automating the application of natural intonation in a voice dialogue system. This system primarily consists of three elements: a server, a terminal, and a user.
[0255] First, the server retrieves historical audio feature data from the company's internal database and preprocesses it into a format suitable for machine learning. The audio data is converted into acoustic features such as MFCCs and spectrograms. This conversion extracts important information from the audio waveform, which can then be used to train the AI model.
[0256] Next, the server builds a deep learning model for the dialogue system and trains it using pre-processed data. This model employs a multi-layer neural network architecture and has the ability to generate speech with natural intonation based on the input text. During the training phase, techniques such as cross-validation are used to improve the accuracy of intonation and expression patterns.
[0257] The user inputs arbitrary text through the terminal. This text data is sent from the terminal to the server, which uses a trained model to generate speech data with natural intonation corresponding to the meaning and context of the text.
[0258] The generated audio data is sent back to the terminal and played back by the user. During this process, the user evaluates whether the generated audio sounds natural and provides feedback as needed. This feedback information is sent to the server and used to further improve the model's accuracy. This feedback loop allows the system to continuously improve, resulting in higher quality generated audio.
[0259] For example, if a user types "Tell me the weather for tomorrow," the server can generate a friendly voice response with an emotional intonation, such as "It's supposed to be sunny tomorrow." This allows the user to have a natural, near-human communication experience.
[0260] The following describes the processing flow.
[0261] Step 1:
[0262] The server retrieves previously collected audio data from the database. This data is stored as pairs of text and corresponding audio files.
[0263] Step 2:
[0264] The server performs preprocessing on the audio data. Specifically, it extracts acoustic features from the audio waveform and converts them into formats such as MFCCs and spectrograms. This makes it easier to represent the characteristics of the audio digitally.
[0265] Step 3:
[0266] The server uses the extracted features to train a deep learning model. This model uses manually adjusted speech data as training data to learn the optimal intonation patterns for text.
[0267] Step 4:
[0268] The user uses the terminal to input text they want to interact with. For example, they might type an everyday question.
[0269] Step 5:
[0270] The terminal sends the entered text to the server. The transmitted data contains the text information necessary for speech generation.
[0271] Step 6:
[0272] The server generates speech data with natural intonation based on the received text, using a trained AI model. During this process, it adjusts pitch, rhythm, and emphasis according to the context of the text.
[0273] Step 7:
[0274] The server sends the generated audio data to the terminal. The data sent includes an audio file in a format playable by the user.
[0275] Step 8:
[0276] The terminal plays the generated voice data for the user. The user can check the naturalness and intonation of the voice.
[0277] Step 9:
[0278] The user provides feedback on the played voice. Evaluate whether the voice is appropriate or if corrections are needed.
[0279] Step 10:
[0280] The terminal sends the feedback from the user to the server. This feedback can be used to adjust and improve the model.
[0281] (Example 1)
[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0283] In a voice dialogue system, it is required to generate voices with intonations that are natural and easy for users to relate to. However, in conventional technologies, there were problems such as insufficient intonation and low user satisfaction. Furthermore, there was a lack of a mechanism to adequately reflect feedback for evaluating the naturalness of the generated voice and improving its quality.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0285] In this invention, the server includes means for acquiring acoustic feature data from a data storage device and performing conversion, means for learning appropriate intonation patterns for language information using a multi-layer learning model, and means for generating acoustic data with natural intonation using a learned model based on the language information input by the user. This enables the provision of natural and user-friendly voices to the user and allows for continuous improvement of the model through feedback.
[0286] The "data storage device" is a device that accumulates acoustic feature data and enables data acquisition as needed.
[0287] The "acoustic feature data" is data obtained by analyzing voice information and numerically represented in forms such as mel-frequency cepstral coefficients and spectrograms.
[0288] The "means for performing conversion" is a process or device that converts the acquired acoustic data into a form that can be processed by a machine learning model.
[0289] The "multi-layer learning model" is a neural network with multiple layers and is a model for learning various intonation patterns for language information.
[0290] The "language information" is data related to the text spoken as voice and its context.
[0291] The "appropriate intonation pattern" is a pattern indicating appropriately adjusted pitch, intensity, and rhythm to achieve natural and user-friendly conversations.
[0292] The "learned model" is a deep learning model trained in advance using data and adjusted to handle specific tasks.
[0293] "Natural intonation" refers to intonation that is similar to the sounds humans actually make in real conversations, and is intended to provide a natural conversational experience.
[0294] "Audio data" refers to digital audio information that can be output as sound.
[0295] "Evaluation information" refers to data that shows the content of feedback and evaluations that users provide to the generated audio.
[0296] "Means of tuning a model" refers to the process of making modifications or updates to a machine learning model based on evaluation information in order to improve its performance.
[0297] This invention includes a process in which a server, a terminal, and a user cooperate to generate speech with natural intonation, as part of a voice dialogue system.
[0298] The server accesses a data storage device that stores acoustic feature data and retrieves the necessary data. This data is converted into formats such as MFCC (Mel-Frequency Cepstral Coefficients) and spectrograms, which serve to extract important features from the speech. The conversion is performed using dedicated speech processing software, and the data is preprocessed into a format suitable for machine learning. LibROSA, a Python-based library, is often used for this purpose.
[0299] Subsequently, the server uses deep learning frameworks such as TensorFlow and PyTorch to build a multi-layer learning model. This model is trained using pre-collected acoustic data with the goal of learning appropriate intonation patterns. Natural language processing techniques are also used in conjunction with intonation learning to take into account the meaning and sentiment of the text.
[0300] The user inputs text into the system through a terminal device. The input text is processed by the terminal and sent to the server as text data. An example of the text is the prompt sentence "Tell me tomorrow's weather". This text is converted by a pre-trained deep learning model into voice data with natural intonation. As a result, the user can obtain a natural and immersive response similar to human conversation.
[0301] The generated voice data is sent back to the terminal and played for the user through a speaker or headphones. Through this playback, the user can evaluate the intonation and naturalness of the generated voice and provide feedback. The feedback is input through the user interface and sent to the server. This information is used to adjust and improve the model, contributing to the quality improvement of the next voice generation process.
[0302] The flow of specific processing in Example 1 will be described using FIG. 11.
[0303] Step 1:
[0304] The server acquires acoustic feature data from the data storage device. In this step, voice recording data is retrieved from the database and output in a voice file format (e.g., WAV).
[0305] Step 2:
[0306] The server preprocesses the acquired acoustic data. In this process, a voice processing library such as LibROSA is used to convert the voice recording data into MFCC or spectrogram. As a result, acoustic features as numerical data that can be used in a machine learning model are output.
[0307] Step 3:
[0308] The server trains a multi-layer learning model using pre-processed data. At this stage, training is performed using TensorFlow or PyTorch, and the output model learns appropriate intonation patterns that take into account pitch and rhythm.
[0309] Step 4:
[0310] The user enters text into the system via the terminal. The entered text is presented as an example prompt, "Tell me tomorrow's weather," and sent to the server as text data.
[0311] Step 5:
[0312] The server inputs the received text data into a trained model and generates audio data with natural intonation. Specifically, it analyzes the text and outputs audio files (e.g., WAV) that reflect the emotion and context.
[0313] Step 6:
[0314] The device receives the generated audio data and plays it for the user via its playback function. In this step, the audio is output and played back through the speaker or headphones.
[0315] Step 7:
[0316] Users evaluate the naturalness and intonation of the played audio and provide feedback. This feedback is entered as an evaluation statement and sent to the server as data for improvement.
[0317] Step 8:
[0318] The server adjusts the model based on the feedback it receives. This feedback is used in the next speech generation process, where data calculations are performed to improve the model's accuracy and produce a more effective model.
[0319] (Application Example 1)
[0320] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0321] Conventional voice dialogue systems suffer from problems such as mechanical and unnatural intonation, making effective communication with users difficult. Furthermore, while there is a demand for natural voice dialogue in home electronic devices, current technology has its limitations. Therefore, there is a need for a means to achieve more natural and human-like voice dialogue.
[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0323] In this invention, the server includes means for acquiring voice feature information from an information storage device and preprocessing this information, means for learning the optimal speech pattern for a sentence using a deep learning model, and means for generating voice information with natural speech based on the sentence input by the user using the trained model. This makes it possible for home electronic devices to provide more natural and effective conversational functions.
[0324] An "information storage device" is a device used to store data such as audio and text, and is used for data retrieval and preprocessing as needed.
[0325] "Voice feature information" refers to the basic attributes and patterns extracted from speech data, and in particular, includes features that contribute to natural intonation.
[0326] "Preprocessing" refers to the process of converting acquired data into a format suitable for analysis and learning, and includes data normalization and feature extraction.
[0327] A "deep learning model" is a model that uses a neural network consisting of multiple layers to learn complex patterns in data.
[0328] The "optimal vocalization pattern" refers to the tone and intonation of the voice that is suitable for achieving a natural and human-like intonation in response to the input text.
[0329] A "user" refers to a person who interacts with this system and utilizes its functions using voice input or text input.
[0330] "Text" refers to the text data to be processed, including the content input for speech generation.
[0331] A "trained model" refers to a model that has been trained using deep learning to specialize in specific patterns or tasks.
[0332] "Voice information with added pronunciation" refers to audio data generated based on text, using natural intonation.
[0333] A "household machine" is an automated machine used within the home, designed to provide a variety of functions, including conversational capabilities.
[0334] "Dialogue functionality" refers to a function that enables humans and machines to exchange information and communicate through voice and text.
[0335] This invention relates to a system for generating natural intonation of speech data. This system is based on three main components: a server, a terminal, and a user.
[0336] The server retrieves voice feature information from its data storage device. This includes previously collected audio data, which is preprocessed to make it suitable for training deep learning models. Specifically, it converts audio waveforms into acoustic features such as MFCCs and spectrograms. This conversion process utilizes servers equipped with NVIDIA GPU devices and TensorFlow software.
[0337] Next, the server constructs a deep learning model, such as a recurrent neural network (RNN) or its variant, the LSTM (Long Short-Term Memory) model, and trains it using pre-processed data. The trained model has the ability to add natural intonation to text and generates speech based on the text data entered by the user.
[0338] Users access the system through a device that allows for voice and text-based interaction. This device can be a smartphone, tablet, or even a robot as a household appliance. Text data provided by the user is sent to the server, and the natural-sounding voice generated by the server is then provided back to the user through the device.
[0339] A concrete example would be a scenario where a household appliance gently informs the user, "Today's weather is sunny," while they are getting ready in the morning, and then prompts them to confirm their daily schedule. An example of a prompt corresponding to this specific example would be, "Generate a response to the user's question with natural intonation."
[0340] This allows users to experience more human-like communication through home electronic devices.
[0341] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0342] Step 1:
[0343] The server retrieves voice feature information from the information storage device. This information includes past audio data and is in a format that allows mapping to text. The data is obtained from the original audio file and used as input for the next preprocessing step.
[0344] Step 2:
[0345] The server preprocesses the acquired audio data. This preprocessing includes conversion to MFCCs and spectrograms. This allows for the extraction of temporal and frequency features of the audio, resulting in high-dimensional audio features. The output is provided as a vector of acoustic features.
[0346] Step 3:
[0347] The server uses pre-processed speech features as input to build and train a deep learning model. Specifically, it uses a recurrent neural network (RNN) or LSTM model to learn the relationship between speech and text. In this step, the model's weights are updated to acquire the optimal speech pattern. Once training is complete, the model is ready to generate speech.
[0348] Step 4:
[0349] The user inputs text through their device and sends that data to the server. The input text data is provided as a sentence to which intonation should be added.
[0350] Step 5:
[0351] The server receives text sent by the user and generates speech using a trained model. The generation process converts the input text into speech data with optimal intonation. The output speech is then sent to the terminal as a digital audio file.
[0352] Step 6:
[0353] The device receives audio data sent from the server and plays it back to the user. This allows the user to hear audio with natural intonation.
[0354] Step 7:
[0355] Users provide feedback on the generated audio. This feedback concerns the naturalness and intelligibility of the audio and is sent to the server and recorded in a database.
[0356] Step 8:
[0357] The server readjusts the model based on user feedback. It analyzes the feedback and uses cross-validation to improve the model's accuracy. This step improves the overall system quality.
[0358] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0359] This invention provides a system for achieving natural intonation and user-centric dialogue in a voice dialogue system. This system consists of four elements: a server, a terminal, a user, and an emotion engine.
[0360] First, the server retrieves speech feature data from the database and preprocesses it. The speech data is processed into acoustic features and converted into a format suitable for training with a deep learning model. Intonation patterns are then learned based on this.
[0361] Next, during the training process using a deep learning model, the server learns the optimal intonation pattern for the user's input text. An emotion engine is incorporated, which can recognize emotions from the user's voice input and text. This engine uses emotion analysis technology to analyze keywords in the text and speech parameters to identify the user's emotional state.
[0362] When a user interacts using a device, the device sends data to the server in response to the user's input (text or voice). The data received by the server is then generated as speech data with natural intonation using a trained model. This generation process also takes into account the user's emotional state, and speech adjustments are made to suit that emotion.
[0363] The generated audio is sent to the device and played back to the user. At this time, the intonation adjusted by the emotion engine provides the user with a more natural and personalized conversational experience. Through this audio playback, the emotional nuances of the speech are more easily conveyed to the user.
[0364] Furthermore, user feedback and evaluations of the voice are sent from the device to the server. This feedback is shared with the emotion engine and used to improve the accuracy of the model and the engine. For example, if a user inputs "I'm feeling a little tired today," the emotion engine recognizes "fatigue," and the server generates a response with an intonation that takes this into consideration. This allows the user to experience more friendly and comfortable communication.
[0365] Thus, the present invention is a system that uses an emotion engine to realize natural voice dialogue optimized for the user.
[0366] The following describes the processing flow.
[0367] Step 1:
[0368] The server retrieves past audio data and its associated emotion labels from the database. This data is stored as text, corresponding audio files, and emotion information.
[0369] Step 2:
[0370] The server extracts acoustic features from the audio data and trains a deep learning model with emotion labels. The model learns the relationship between intonation patterns and emotional states in relation to speech, creating a foundation for natural speech generation.
[0371] Step 3:
[0372] The user inputs the text they want to speak through the device. Sometimes, the user's voice is also input at the same time.
[0373] Step 4:
[0374] The terminal sends user input text and voice to the server. An emotion engine also operates at this point to analyze the emotion in the voice.
[0375] Step 5:
[0376] The server uses an emotion engine to recognize emotions from the user's voice and adjusts the intonation generated by the deep learning model based on that emotion information.
[0377] Step 6:
[0378] The server generates audio data with the optimal intonation for the user's text. The generated audio takes into account the user's emotional state.
[0379] Step 7:
[0380] The server sends the generated voice data to the terminal. The voice, adjusted with emotional information, is more natural and approachable.
[0381] Step 8:
[0382] The device plays the generated audio to the user, who then receives feedback from the system.
[0383] Step 9:
[0384] Users provide feedback on the played audio, evaluating whether it is appropriate and what needs improvement.
[0385] Step 10:
[0386] The device sends user feedback to the server. This feedback is used for the continuous improvement of the emotion engine and deep learning models.
[0387] (Example 2)
[0388] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0389] In voice dialogue systems, generating voice responses with natural intonation that takes into account the user's emotions has been difficult with conventional technologies. Furthermore, there is a need for a system that effectively utilizes user feedback to improve model accuracy.
[0390] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0391] In this invention, the server includes means for preprocessing voice information obtained from a database, means for learning the optimal intonation pattern for the information using a machine learning algorithm, means for generating voice information with natural intonation using a trained model based on information input by the user, means for identifying the emotional state from the input information using sentiment analysis technology, and means for modifying the model based on the evaluation and emotional state identification results. This enables the generation of natural voice responses that correspond to the user's emotional state and continuous model improvement.
[0392] A "database" is a storage or management system that systematically organizes information to facilitate searching and access.
[0393] "Vocal information" refers to data related to human speech, and includes acoustic features, intonation, and parameters indicating emotion.
[0394] "Preprocessing" is the stage of shaping or transforming raw data into a format suitable for subsequent processing, and involves denoising data and extracting features.
[0395] A "machine learning algorithm" is a set of logic and procedures that allows a computer to learn patterns from data and perform predictions and classifications.
[0396] An "intonation pattern" is a structure that indicates the placement of intonation and accent in speech, and is important for achieving natural speech.
[0397] A "trained model" is an algorithm that has undergone an initial training process and incorporates knowledge based on diverse data.
[0398] "Sentiment analysis technology" is a technology that automatically identifies and classifies a person's emotions expressed in text or audio.
[0399] "Evaluation" refers to the act of users reviewing the results generated by a system and providing feedback based on quality and satisfaction.
[0400] "Modification" refers to making adjustments or changes to a system or algorithm in order to improve its initial state or its state after changes.
[0401] As an embodiment of this invention, details of a dialogue system using voice information are provided. The system consists of a server, terminals, users, and a data communication network connecting them.
[0402] The server acts as a central processing unit, handling audio information processing. Specifically, it retrieves audio information from the database and performs preprocessing. This preprocessing involves using audio processing libraries (e.g., Librosa) to remove noise and extract acoustic features. These features are then processed by deep learning models (e.g., RNNs and Transformers) to learn intonation patterns. The server also incorporates sentiment analysis technology, which uses natural language processing tools (e.g., NLTK and TextBlob) to identify emotions.
[0403] The user inputs voice or text via a terminal acting as an interactive device. The user's input is transmitted to the server through the terminal. The terminal then forwards the received voice data or text information directly to the server.
[0404] The device receives voice information generated from the server. This voice information is generated with intonation that takes the user's emotions into consideration, providing a more natural conversational experience. The user can provide feedback on this played voice information through the device.
[0405] The server accumulates user evaluations and feedback, and continuously improves the generated AI model. Cross-validation technology can be used to improve the model's accuracy.
[0406] As a concrete example, here is an example of a prompt: "Based on the text received from the user, generate a voice response with natural intonation that takes into account the emotional state." This prompt enables the system to engage in highly engaging conversations with the user, providing efficient and satisfying communication.
[0407] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0408] Step 1:
[0409] The server retrieves audio information from the database. It receives audio samples and associated metadata as input. This audio information contains various intonation patterns and is the target of analysis. The server uses an audio processing library to extract fundamental frequencies and spectral information and convert them into acoustic features. The output provides acoustic features for learning intonation.
[0410] Step 2:
[0411] The server inputs the acquired acoustic features into a deep learning model to learn the optimal intonation pattern. Machine learning algorithms are used here. During this process, the model acquires the ability to recognize patterns in information and identify intonation patterns. The learned intonation pattern is provided as output.
[0412] Step 3:
[0413] The user inputs data via voice or text through the device. This input includes data that reflects the user's intentions and emotions. The device sends this data to the server in preparation for the next processing step.
[0414] Step 4:
[0415] The server uses sentiment analysis technology to identify the user's emotional state based on input data received from the user. Input can be text or audio, along with a currently learned model. Here, natural language processing tools are used to analyze keywords and phrases to determine the user's emotions. The identified emotional state is provided as output.
[0416] Step 5:
[0417] The server uses a trained model to adjust intonation according to the identified emotional state and generate a voice response. The input consists of trained intonation patterns and the user's emotion. The output is voice data with a natural intonation that matches the user's emotion.
[0418] Step 6:
[0419] The device receives the generated audio and plays it for the user. It handles audio data sent from the server as input. During playback, it ensures that the intonation is natural and aligns with the user's emotions. The final audio experience delivered to the user is then produced as output.
[0420] Step 7:
[0421] Users provide feedback on voice responses. As input, they enter their evaluation or comments on the generated voice into the device. The device sends this to the server, which helps improve the model. As output, the feedback is accumulated and used to generate the next response.
[0422] (Application Example 2)
[0423] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0424] Conventional voice dialogue systems have struggled to provide natural intonation that fully considers the user's emotions, making it difficult to achieve emotionally resonant dialogue. This problem makes it challenging to provide friendly and comfortable communication for users.
[0425] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0426] In this invention, the server includes means for acquiring speech feature information from a data set and preprocessing it, means for learning the optimal intonation pattern for the recording using a machine learning system, and means for analyzing the user's emotions and reflecting emotion-based intonation in the speech information. This makes it possible to provide emotion-adaptive speech dialogue with natural intonation.
[0427] A "data set" is a collection of data that is gathered and managed for a specific purpose.
[0428] "Audio feature information" refers to information containing features such as frequency, intensity, and pitch extracted from audio data.
[0429] "Preprocessing" is the process of transforming and processing data to make it suitable for analysis and learning.
[0430] A "machine learning system" is a system that uses algorithms to learn rules and patterns from data and perform predictions and classifications.
[0431] "Records" are data or documents that are stored and protected for later reference.
[0432] An "intonation pattern" is a pattern that shows the combination and arrangement of intonation in the sounds of a language.
[0433] A "user" is a person who uses a system or device.
[0434] "Emotional analysis" is the process of extracting emotional states from audio or text as numerical values or labels.
[0435] "Audio information" refers to all information transmitted as sound, including audio signal data.
[0436] "Emotionally adaptive voice interaction" refers to voice interaction that recognizes the user's emotions and engages in natural conversation tailored to those emotions.
[0437] The system implementing this invention consists of three elements: a server, a terminal, and a user. The server acquires speech feature information from a data set and preprocesses that data. The preprocessed data is used by a machine learning system to learn the optimal intonation pattern for recording. In particular, TensorFlow / Keras is used as a deep learning framework to learn natural intonation. Speech processing tools such as PyDub are utilized for processing the speech data.
[0438] The server then generates speech information with natural intonation using a trained system based on the user's input. The terminal receives the user's input and sends that information to the server. The server analyzes the user's emotions using an emotion recognition engine. Using the natural language processing library NLTK, it determines the user's emotions and generates speech information with the appropriate intonation.
[0439] The generated voice information is sent to the terminal and played back by the user. Through the played voice, emotionally adaptive voice dialogue is realized. This dialogue can be used particularly in home assist robots and improves the quality of everyday communication.
[0440] As a specific example, in a morning scenario, when the robot says, "Good morning. How are you feeling today?", and the user inputs, "I'm feeling a little down," the robot will provide an emotionally adaptive response. In this case, the robot will reply in a gentle tone of voice, "I see. Let's do something to help you relax."
[0441] An example of a prompt is, "Generate a kind and considerate intonation pattern for when a user says, 'Today was tough.'" This sentence is used in a generative AI model to generate responses appropriate to the user's emotions.
[0442] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0443] Step 1:
[0444] The terminal acquires text and voice data entered by the user. This data is sent to the server as input data. The input data is expected to contain the user's intentions and emotions.
[0445] Step 2:
[0446] The server extracts speech feature information from the received speech data and performs natural language processing on the text data as needed. It extracts acoustic features from the speech data and performs emotion recognition on the text data. Based on the input data, it analyzes the features of the speech data and generates a feature vector as an intermediate output.
[0447] Step 3:
[0448] The server uses a deep learning model to generate the optimal intonation pattern from the feature vector. Here, the learned pattern is applied using a generative AI model. The output of this step is the generated intonation pattern.
[0449] Step 4:
[0450] The server uses an emotion recognition engine to analyze the user's emotions and adjust the intonation accordingly. Specifically, it takes the emotion analysis results as input and adjusts the intonation appropriately. It generates speech data with intonation that matches the emotion.
[0451] Step 5:
[0452] The server sends the generated audio data to the terminal. The terminal plays this data back to the user. This provides the user with natural and emotionally appropriate voice responses. The final output is the intoned audio played back to the user.
[0453] Step 6:
[0454] The user evaluates the played audio and inputs their feedback into the device. The device then sends the feedback to the server. This evaluation information is used to further improve the system.
[0455] Step 7:
[0456] Based on the feedback received, the server adjusts the speech generation system and improves the accuracy of the learning model and emotion engine in preparation for the next use. This improves the overall performance of the system.
[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0460] [Third Embodiment]
[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0473] This invention provides a system for automating the application of natural intonation in a voice dialogue system. This system primarily consists of three elements: a server, a terminal, and a user.
[0474] First, the server retrieves historical audio feature data from the company's internal database and preprocesses it into a format suitable for machine learning. The audio data is converted into acoustic features such as MFCCs and spectrograms. This conversion extracts important information from the audio waveform, which can then be used to train the AI model.
[0475] Next, the server builds a deep learning model for the dialogue system and trains it using pre-processed data. This model employs a multi-layer neural network architecture and has the ability to generate speech with natural intonation based on the input text. During the training phase, techniques such as cross-validation are used to improve the accuracy of intonation and expression patterns.
[0476] The user inputs arbitrary text through the terminal. This text data is sent from the terminal to the server, which uses a trained model to generate speech data with natural intonation corresponding to the meaning and context of the text.
[0477] The generated audio data is sent back to the terminal and played back by the user. During this process, the user evaluates whether the generated audio sounds natural and provides feedback as needed. This feedback information is sent to the server and used to further improve the model's accuracy. This feedback loop allows the system to continuously improve, resulting in higher quality generated audio.
[0478] For example, if a user types "Tell me the weather for tomorrow," the server can generate a friendly voice response with an emotional intonation, such as "It's supposed to be sunny tomorrow." This allows the user to have a natural, near-human communication experience.
[0479] The following describes the processing flow.
[0480] Step 1:
[0481] The server retrieves previously collected audio data from the database. This data is stored as pairs of text and corresponding audio files.
[0482] Step 2:
[0483] The server performs preprocessing on the audio data. Specifically, it extracts acoustic features from the audio waveform and converts them into formats such as MFCCs and spectrograms. This makes it easier to represent the characteristics of the audio digitally.
[0484] Step 3:
[0485] The server uses the extracted features to train a deep learning model. This model uses manually adjusted speech data as training data to learn the optimal intonation patterns for text.
[0486] Step 4:
[0487] The user uses the terminal to input text they want to interact with. For example, they might type an everyday question.
[0488] Step 5:
[0489] The terminal sends the entered text to the server. The transmitted data contains the text information necessary for speech generation.
[0490] Step 6:
[0491] The server generates speech data with natural intonation based on the received text, using a trained AI model. During this process, it adjusts pitch, rhythm, and emphasis according to the context of the text.
[0492] Step 7:
[0493] The server sends the generated audio data to the terminal. The data sent includes an audio file in a format playable by the user.
[0494] Step 8:
[0495] The device plays the generated audio data to the user. The user can then check the naturalness and intonation of the audio.
[0496] Step 9:
[0497] Users provide feedback on the played audio, evaluating whether it is appropriate and whether it needs correction.
[0498] Step 10:
[0499] The device sends user feedback to the server. This feedback is used to adjust and improve the model.
[0500] (Example 1)
[0501] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] In voice dialogue systems, there is a need to generate speech with natural and user-friendly intonation. However, conventional technologies have suffered from insufficient intonation, resulting in low user satisfaction. Furthermore, there is a lack of mechanisms to adequately evaluate the naturalness of the generated speech and incorporate feedback to improve its quality.
[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0504] In this invention, the server includes means for acquiring and converting acoustic feature data from a data storage device, means for learning appropriate intonation patterns for linguistic information using a multilayer learning model, and means for generating acoustic data with natural intonation using the trained model based on linguistic information input by the user. This provides users with natural and familiar speech and enables continuous model improvement through feedback.
[0505] A "data storage device" is a device that stores acoustic characteristic data and allows the data to be retrieved as needed.
[0506] "Acoustic feature data" refers to data obtained by analyzing audio information and representing it numerically in forms such as Mel-frequency cepstrum coefficients and spectrograms.
[0507] "Means of conversion" refers to a process or device that converts acquired acoustic data into a format that is easily processed by a machine learning model.
[0508] A "multilayer learning model" is a neural network with multiple layers, designed to learn various intonation patterns for linguistic information.
[0509] "Linguistic information" refers to data about text spoken as audio and its context.
[0510] An "appropriate intonation pattern" is a pattern that shows appropriately adjusted pitch, volume, and rhythm to achieve natural and friendly conversation.
[0511] A "pre-trained model" is a deep learning model that has been trained in advance using data and is tuned to handle a specific task.
[0512] "Natural intonation" refers to intonation that is similar to the sounds humans actually make in real conversations, and is intended to provide a natural conversational experience.
[0513] "Audio data" refers to digital audio information that can be output as sound.
[0514] "Evaluation information" refers to data that shows the content of feedback and evaluations that users provide to the generated audio.
[0515] "Means of tuning a model" refers to the process of making modifications or updates to a machine learning model based on evaluation information in order to improve its performance.
[0516] This invention includes a process in which a server, a terminal, and a user cooperate to generate speech with natural intonation, as part of a voice dialogue system.
[0517] The server accesses a data storage device that stores acoustic feature data and retrieves the necessary data. This data is converted into formats such as MFCC (Mel-Frequency Cepstral Coefficients) and spectrograms, which serve to extract important features from the speech. The conversion is performed using dedicated speech processing software, and the data is preprocessed into a format suitable for machine learning. LibROSA, a Python-based library, is often used for this purpose.
[0518] Subsequently, the server uses deep learning frameworks such as TensorFlow and PyTorch to build a multi-layer learning model. This model is trained using pre-collected acoustic data with the goal of learning appropriate intonation patterns. Natural language processing techniques are also used in conjunction with intonation learning to take into account the meaning and sentiment of the text.
[0519] The user inputs text into the system via a terminal device. The input text is processed by the terminal and sent to the server as text data. An example of text is the prompt "Tell me the weather tomorrow." This text is converted into speech data with natural intonation by a trained deep learning model. As a result, the user can receive a natural and immersive response that closely resembles human dialogue.
[0520] The generated audio data is sent back to the device and played back to the user through speakers or headphones. This playback allows the user to evaluate the intonation and naturalness of the generated audio and provide feedback. This feedback is entered through the user interface and sent to the server. This information is used to adjust and improve the model, contributing to improved quality in subsequent audio generation processes.
[0521] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0522] Step 1:
[0523] The server retrieves acoustic feature data from the data storage device. In this step, audio recording data is retrieved from the database and output as an audio file format (e.g., WAV).
[0524] Step 2:
[0525] The server preprocesses the acquired acoustic data. This process uses an audio processing library such as LibROSA to convert the audio recording data into MFCCs or spectrograms. As a result, it outputs acoustic features as numerical data usable in machine learning models.
[0526] Step 3:
[0527] The server trains a multi-layer learning model using pre-processed data. At this stage, training is performed using TensorFlow or PyTorch, and the output model learns appropriate intonation patterns that take into account pitch and rhythm.
[0528] Step 4:
[0529] The user enters text into the system via the terminal. The entered text is presented as an example prompt, "Tell me tomorrow's weather," and sent to the server as text data.
[0530] Step 5:
[0531] The server inputs the received text data into a trained model and generates audio data with natural intonation. Specifically, it analyzes the text and outputs audio files (e.g., WAV) that reflect the emotion and context.
[0532] Step 6:
[0533] The device receives the generated audio data and plays it for the user via its playback function. In this step, the audio is output and played back through the speaker or headphones.
[0534] Step 7:
[0535] Users evaluate the naturalness and intonation of the played audio and provide feedback. This feedback is entered as an evaluation statement and sent to the server as data for improvement.
[0536] Step 8:
[0537] The server adjusts the model based on the feedback it receives. This feedback is used in the next speech generation process, where data calculations are performed to improve the model's accuracy and produce a more effective model.
[0538] (Application Example 1)
[0539] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0540] Conventional voice dialogue systems suffer from problems such as mechanical and unnatural intonation, making effective communication with users difficult. Furthermore, while there is a demand for natural voice dialogue in home electronic devices, current technology has its limitations. Therefore, there is a need for a means to achieve more natural and human-like voice dialogue.
[0541] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0542] In this invention, the server includes means for acquiring voice feature information from an information storage device and preprocessing this information, means for learning the optimal speech pattern for a sentence using a deep learning model, and means for generating voice information with natural speech based on the sentence input by the user using the trained model. This makes it possible for home electronic devices to provide more natural and effective conversational functions.
[0543] An "information storage device" is a device used to store data such as audio and text, and is used for data retrieval and preprocessing as needed.
[0544] "Voice feature information" refers to the basic attributes and patterns extracted from speech data, and in particular, includes features that contribute to natural intonation.
[0545] "Preprocessing" refers to the process of converting acquired data into a format suitable for analysis and learning, and includes data normalization and feature extraction.
[0546] A "deep learning model" is a model that uses a neural network consisting of multiple layers to learn complex patterns in data.
[0547] The "optimal vocalization pattern" refers to the tone and intonation of the voice that is suitable for achieving a natural and human-like intonation in response to the input text.
[0548] A "user" refers to a person who interacts with this system and utilizes its functions using voice input or text input.
[0549] "Text" refers to the text data to be processed, including the content input for speech generation.
[0550] A "trained model" refers to a model that has been trained using deep learning to specialize in specific patterns or tasks.
[0551] "Voice information with added pronunciation" refers to audio data generated based on text, using natural intonation.
[0552] A "household machine" is an automated machine used within the home, designed to provide a variety of functions, including conversational capabilities.
[0553] "Dialogue functionality" refers to a function that enables humans and machines to exchange information and communicate through voice and text.
[0554] This invention relates to a system for generating natural intonation of speech data. This system is based on three main components: a server, a terminal, and a user.
[0555] The server retrieves voice feature information from its data storage device. This includes previously collected audio data, which is preprocessed to make it suitable for training deep learning models. Specifically, it converts audio waveforms into acoustic features such as MFCCs and spectrograms. This conversion process utilizes servers equipped with NVIDIA GPU devices and TensorFlow software.
[0556] Next, the server constructs a deep learning model, such as a recurrent neural network (RNN) or its variant, the LSTM (Long Short-Term Memory) model, and trains it using pre-processed data. The trained model has the ability to add natural intonation to text and generates speech based on the text data entered by the user.
[0557] Users access the system through a device that allows for voice and text-based interaction. This device can be a smartphone, tablet, or even a robot as a household appliance. Text data provided by the user is sent to the server, and the natural-sounding voice generated by the server is then provided back to the user through the device.
[0558] A concrete example would be a scenario where a household appliance gently informs the user, "Today's weather is sunny," while they are getting ready in the morning, and then prompts them to confirm their daily schedule. An example of a prompt corresponding to this specific example would be, "Generate a response to the user's question with natural intonation."
[0559] This allows users to experience more human-like communication through home electronic devices.
[0560] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0561] Step 1:
[0562] The server retrieves voice feature information from the information storage device. This information includes past audio data and is in a format that allows mapping to text. The data is obtained from the original audio file and used as input for the next preprocessing step.
[0563] Step 2:
[0564] The server preprocesses the acquired audio data. This preprocessing includes conversion to MFCCs and spectrograms. This allows for the extraction of temporal and frequency features of the audio, resulting in high-dimensional audio features. The output is provided as a vector of acoustic features.
[0565] Step 3:
[0566] The server uses pre-processed speech features as input to build and train a deep learning model. Specifically, it uses a recurrent neural network (RNN) or LSTM model to learn the relationship between speech and text. In this step, the model's weights are updated to acquire the optimal speech pattern. Once training is complete, the model is ready to generate speech.
[0567] Step 4:
[0568] The user inputs text through their device and sends that data to the server. The input text data is provided as a sentence to which intonation should be added.
[0569] Step 5:
[0570] The server receives text sent by the user and generates speech using a trained model. The generation process converts the input text into speech data with optimal intonation. The output speech is then sent to the terminal as a digital audio file.
[0571] Step 6:
[0572] The device receives audio data sent from the server and plays it back to the user. This allows the user to hear audio with natural intonation.
[0573] Step 7:
[0574] Users provide feedback on the generated audio. This feedback concerns the naturalness and intelligibility of the audio and is sent to the server and recorded in a database.
[0575] Step 8:
[0576] The server readjusts the model based on user feedback. It analyzes the feedback and uses cross-validation to improve the model's accuracy. This step improves the overall system quality.
[0577] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0578] This invention provides a system for achieving natural intonation and user-centric dialogue in a voice dialogue system. This system consists of four elements: a server, a terminal, a user, and an emotion engine.
[0579] First, the server retrieves speech feature data from the database and preprocesses it. The speech data is processed into acoustic features and converted into a format suitable for training with a deep learning model. Intonation patterns are then learned based on this.
[0580] Next, during the training process using a deep learning model, the server learns the optimal intonation pattern for the user's input text. An emotion engine is incorporated, which can recognize emotions from the user's voice input and text. This engine uses emotion analysis technology to analyze keywords in the text and speech parameters to identify the user's emotional state.
[0581] When a user interacts using a device, the device sends data to the server in response to the user's input (text or voice). The data received by the server is then generated as speech data with natural intonation using a trained model. This generation process also takes into account the user's emotional state, and speech adjustments are made to suit that emotion.
[0582] The generated audio is sent to the device and played back to the user. At this time, the intonation adjusted by the emotion engine provides the user with a more natural and personalized conversational experience. Through this audio playback, the emotional nuances of the speech are more easily conveyed to the user.
[0583] Furthermore, user feedback and evaluations of the voice are sent from the device to the server. This feedback is shared with the emotion engine and used to improve the accuracy of the model and the engine. For example, if a user inputs "I'm feeling a little tired today," the emotion engine recognizes "fatigue," and the server generates a response with an intonation that takes this into consideration. This allows the user to experience more friendly and comfortable communication.
[0584] Thus, the present invention is a system that uses an emotion engine to realize natural voice dialogue optimized for the user.
[0585] The following describes the processing flow.
[0586] Step 1:
[0587] The server retrieves past audio data and its associated emotion labels from the database. This data is stored as text, corresponding audio files, and emotion information.
[0588] Step 2:
[0589] The server extracts acoustic features from the audio data and trains a deep learning model with emotion labels. The model learns the relationship between intonation patterns and emotional states in relation to speech, creating a foundation for natural speech generation.
[0590] Step 3:
[0591] The user inputs the text they want to speak through the device. Sometimes, the user's voice is also input at the same time.
[0592] Step 4:
[0593] The terminal sends user input text and voice to the server. An emotion engine also operates at this point to analyze the emotion in the voice.
[0594] Step 5:
[0595] The server uses an emotion engine to recognize emotions from the user's voice and adjusts the intonation generated by the deep learning model based on that emotion information.
[0596] Step 6:
[0597] The server generates audio data with the optimal intonation for the user's text. The generated audio takes into account the user's emotional state.
[0598] Step 7:
[0599] The server sends the generated voice data to the terminal. The voice, adjusted with emotional information, is more natural and approachable.
[0600] Step 8:
[0601] The device plays the generated audio to the user, who then receives feedback from the system.
[0602] Step 9:
[0603] Users provide feedback on the played audio, evaluating whether it is appropriate and what needs improvement.
[0604] Step 10:
[0605] The device sends user feedback to the server. This feedback is used for the continuous improvement of the emotion engine and deep learning models.
[0606] (Example 2)
[0607] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0608] In voice dialogue systems, generating voice responses with natural intonation that takes into account the user's emotions has been difficult with conventional technologies. Furthermore, there is a need for a system that effectively utilizes user feedback to improve model accuracy.
[0609] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0610] In this invention, the server includes means for preprocessing voice information obtained from a database, means for learning the optimal intonation pattern for the information using a machine learning algorithm, means for generating voice information with natural intonation using a trained model based on information input by the user, means for identifying the emotional state from the input information using sentiment analysis technology, and means for modifying the model based on the evaluation and emotional state identification results. This enables the generation of natural voice responses that correspond to the user's emotional state and continuous model improvement.
[0611] A "database" is a storage or management system that systematically organizes information to facilitate searching and access.
[0612] "Vocal information" refers to data related to human speech, and includes acoustic features, intonation, and parameters indicating emotion.
[0613] "Preprocessing" is the stage of shaping or transforming raw data into a format suitable for subsequent processing, and involves denoising data and extracting features.
[0614] A "machine learning algorithm" is a set of logic and procedures that allows a computer to learn patterns from data and perform predictions and classifications.
[0615] An "intonation pattern" is a structure that indicates the placement of intonation and accent in speech, and is important for achieving natural speech.
[0616] A "trained model" is an algorithm that has undergone an initial training process and incorporates knowledge based on diverse data.
[0617] "Sentiment analysis technology" is a technology that automatically identifies and classifies a person's emotions expressed in text or audio.
[0618] "Evaluation" refers to the act of users reviewing the results generated by a system and providing feedback based on quality and satisfaction.
[0619] "Modification" refers to making adjustments or changes to a system or algorithm in order to improve its initial state or its state after changes.
[0620] As an embodiment of this invention, details of a dialogue system using voice information are provided. The system consists of a server, terminals, users, and a data communication network connecting them.
[0621] The server acts as a central processing unit, handling audio information processing. Specifically, it retrieves audio information from the database and performs preprocessing. This preprocessing involves using audio processing libraries (e.g., Librosa) to remove noise and extract acoustic features. These features are then processed by deep learning models (e.g., RNNs and Transformers) to learn intonation patterns. The server also incorporates sentiment analysis technology, which uses natural language processing tools (e.g., NLTK and TextBlob) to identify emotions.
[0622] The user inputs voice or text via a terminal acting as an interactive device. The user's input is transmitted to the server through the terminal. The terminal then forwards the received voice data or text information directly to the server.
[0623] The device receives voice information generated from the server. This voice information is generated with intonation that takes the user's emotions into consideration, providing a more natural conversational experience. The user can provide feedback on this played voice information through the device.
[0624] The server accumulates user evaluations and feedback, and continuously improves the generated AI model. Cross-validation technology can be used to improve the model's accuracy.
[0625] As a concrete example, here is an example of a prompt: "Based on the text received from the user, generate a voice response with natural intonation that takes into account the emotional state." This prompt enables the system to engage in highly engaging conversations with the user, providing efficient and satisfying communication.
[0626] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0627] Step 1:
[0628] The server retrieves audio information from the database. It receives audio samples and associated metadata as input. This audio information contains various intonation patterns and is the target of analysis. The server uses an audio processing library to extract fundamental frequencies and spectral information and convert them into acoustic features. The output provides acoustic features for learning intonation.
[0629] Step 2:
[0630] The server inputs the acquired acoustic features into a deep learning model to learn the optimal intonation pattern. Machine learning algorithms are used here. During this process, the model acquires the ability to recognize patterns in information and identify intonation patterns. The learned intonation pattern is provided as output.
[0631] Step 3:
[0632] The user inputs data via voice or text through the device. This input includes data that reflects the user's intentions and emotions. The device sends this data to the server in preparation for the next processing step.
[0633] Step 4:
[0634] The server uses sentiment analysis technology to identify the user's emotional state based on input data received from the user. Input can be text or audio, along with a currently learned model. Here, natural language processing tools are used to analyze keywords and phrases to determine the user's emotions. The identified emotional state is provided as output.
[0635] Step 5:
[0636] The server uses a trained model to adjust intonation according to the identified emotional state and generate a voice response. The input consists of trained intonation patterns and the user's emotion. The output is voice data with a natural intonation that matches the user's emotion.
[0637] Step 6:
[0638] The device receives the generated audio and plays it for the user. It handles audio data sent from the server as input. During playback, it ensures that the intonation is natural and aligns with the user's emotions. The final audio experience delivered to the user is then produced as output.
[0639] Step 7:
[0640] Users provide feedback on voice responses. As input, they enter their evaluation or comments on the generated voice into the device. The device sends this to the server, which helps improve the model. As output, the feedback is accumulated and used to generate the next response.
[0641] (Application Example 2)
[0642] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0643] Conventional voice dialogue systems have struggled to provide natural intonation that fully considers the user's emotions, making it difficult to achieve emotionally resonant dialogue. This problem makes it challenging to provide friendly and comfortable communication for users.
[0644] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0645] In this invention, the server includes means for acquiring speech feature information from a data set and preprocessing it, means for learning the optimal intonation pattern for the recording using a machine learning system, and means for analyzing the user's emotions and reflecting emotion-based intonation in the speech information. This makes it possible to provide emotion-adaptive speech dialogue with natural intonation.
[0646] A "data set" is a collection of data that is gathered and managed for a specific purpose.
[0647] "Audio feature information" refers to information containing features such as frequency, intensity, and pitch extracted from audio data.
[0648] "Preprocessing" is the process of transforming and processing data to make it suitable for analysis and learning.
[0649] A "machine learning system" is a system that uses algorithms to learn rules and patterns from data and perform predictions and classifications.
[0650] "Records" are data or documents that are stored and protected for later reference.
[0651] An "intonation pattern" is a pattern that shows the combination and arrangement of intonation in the sounds of a language.
[0652] A "user" is a person who uses a system or device.
[0653] "Emotional analysis" is the process of extracting emotional states from audio or text as numerical values or labels.
[0654] "Audio information" refers to all information transmitted as sound, including audio signal data.
[0655] "Emotionally adaptive voice interaction" refers to voice interaction that recognizes the user's emotions and engages in natural conversation tailored to those emotions.
[0656] The system implementing this invention consists of three elements: a server, a terminal, and a user. The server acquires speech feature information from a data set and preprocesses that data. The preprocessed data is used by a machine learning system to learn the optimal intonation pattern for recording. In particular, TensorFlow / Keras is used as a deep learning framework to learn natural intonation. Speech processing tools such as PyDub are utilized for processing the speech data.
[0657] The server then generates speech information with natural intonation using a trained system based on the user's input. The terminal receives the user's input and sends that information to the server. The server analyzes the user's emotions using an emotion recognition engine. Using the natural language processing library NLTK, it determines the user's emotions and generates speech information with the appropriate intonation.
[0658] The generated voice information is sent to the terminal and played back by the user. Through the played voice, emotionally adaptive voice dialogue is realized. This dialogue can be used particularly in home assist robots and improves the quality of everyday communication.
[0659] As a specific example, in a morning scenario, when the robot says, "Good morning. How are you feeling today?", and the user inputs, "I'm feeling a little down," the robot will provide an emotionally adaptive response. In this case, the robot will reply in a gentle tone of voice, "I see. Let's do something to help you relax."
[0660] An example of a prompt is, "Generate a kind and considerate intonation pattern for when a user says, 'Today was tough.'" This sentence is used in a generative AI model to generate responses appropriate to the user's emotions.
[0661] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0662] Step 1:
[0663] The terminal acquires text and voice data entered by the user. This data is sent to the server as input data. The input data is expected to contain the user's intentions and emotions.
[0664] Step 2:
[0665] The server extracts speech feature information from the received speech data and performs natural language processing on the text data as needed. It extracts acoustic features from the speech data and performs emotion recognition on the text data. Based on the input data, it analyzes the features of the speech data and generates a feature vector as an intermediate output.
[0666] Step 3:
[0667] The server uses a deep learning model to generate the optimal intonation pattern from the feature vector. Here, the learned pattern is applied using a generative AI model. The output of this step is the generated intonation pattern.
[0668] Step 4:
[0669] The server uses an emotion recognition engine to analyze the user's emotions and adjust the intonation accordingly. Specifically, it takes the emotion analysis results as input and adjusts the intonation appropriately. It generates speech data with intonation that matches the emotion.
[0670] Step 5:
[0671] The server sends the generated audio data to the terminal. The terminal plays this data back to the user. This provides the user with natural and emotionally appropriate voice responses. The final output is the intoned audio played back to the user.
[0672] Step 6:
[0673] The user evaluates the played audio and inputs their feedback into the device. The device then sends the feedback to the server. This evaluation information is used to further improve the system.
[0674] Step 7:
[0675] Based on the feedback received, the server adjusts the speech generation system and improves the accuracy of the learning model and emotion engine in preparation for the next use. This improves the overall performance of the system.
[0676] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0678] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0679] [Fourth Embodiment]
[0680] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0681] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0682] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0683] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0684] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0685] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0686] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0687] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0688] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0691] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] This invention provides a system for automating the application of natural intonation in a voice dialogue system. This system primarily consists of three elements: a server, a terminal, and a user.
[0694] First, the server retrieves historical audio feature data from the company's internal database and preprocesses it into a format suitable for machine learning. The audio data is converted into acoustic features such as MFCCs and spectrograms. This conversion extracts important information from the audio waveform, which can then be used to train the AI model.
[0695] Next, the server builds a deep learning model for the dialogue system and trains it using pre-processed data. This model employs a multi-layer neural network architecture and has the ability to generate speech with natural intonation based on the input text. During the training phase, techniques such as cross-validation are used to improve the accuracy of intonation and expression patterns.
[0696] The user inputs arbitrary text through the terminal. This text data is sent from the terminal to the server, which uses a trained model to generate speech data with natural intonation corresponding to the meaning and context of the text.
[0697] The generated audio data is sent back to the terminal and played back by the user. During this process, the user evaluates whether the generated audio sounds natural and provides feedback as needed. This feedback information is sent to the server and used to further improve the model's accuracy. This feedback loop allows the system to continuously improve, resulting in higher quality generated audio.
[0698] For example, if a user types "Tell me the weather for tomorrow," the server can generate a friendly voice response with an emotional intonation, such as "It's supposed to be sunny tomorrow." This allows the user to have a natural, near-human communication experience.
[0699] The following describes the processing flow.
[0700] Step 1:
[0701] The server retrieves previously collected audio data from the database. This data is stored as pairs of text and corresponding audio files.
[0702] Step 2:
[0703] The server performs preprocessing on the audio data. Specifically, it extracts acoustic features from the audio waveform and converts them into formats such as MFCCs and spectrograms. This makes it easier to represent the characteristics of the audio digitally.
[0704] Step 3:
[0705] The server uses the extracted features to train a deep learning model. This model uses manually adjusted speech data as training data to learn the optimal intonation patterns for text.
[0706] Step 4:
[0707] The user uses the terminal to input text they want to interact with. For example, they might type an everyday question.
[0708] Step 5:
[0709] The terminal sends the entered text to the server. The transmitted data contains the text information necessary for speech generation.
[0710] Step 6:
[0711] The server generates speech data with natural intonation based on the received text, using a trained AI model. During this process, it adjusts pitch, rhythm, and emphasis according to the context of the text.
[0712] Step 7:
[0713] The server sends the generated audio data to the terminal. The data sent includes an audio file in a format playable by the user.
[0714] Step 8:
[0715] The device plays the generated audio data to the user. The user can then check the naturalness and intonation of the audio.
[0716] Step 9:
[0717] Users provide feedback on the played audio, evaluating whether it is appropriate and whether it needs correction.
[0718] Step 10:
[0719] The device sends user feedback to the server. This feedback is used to adjust and improve the model.
[0720] (Example 1)
[0721] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0722] In voice dialogue systems, there is a need to generate speech with natural and user-friendly intonation. However, conventional technologies have suffered from insufficient intonation, resulting in low user satisfaction. Furthermore, there is a lack of mechanisms to adequately evaluate the naturalness of the generated speech and incorporate feedback to improve its quality.
[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0724] In this invention, the server includes means for acquiring and converting acoustic feature data from a data storage device, means for learning appropriate intonation patterns for linguistic information using a multilayer learning model, and means for generating acoustic data with natural intonation using the trained model based on linguistic information input by the user. This provides users with natural and familiar speech and enables continuous model improvement through feedback.
[0725] A "data storage device" is a device that stores acoustic characteristic data and allows the data to be retrieved as needed.
[0726] "Acoustic feature data" refers to data obtained by analyzing audio information and representing it numerically in forms such as Mel-frequency cepstrum coefficients and spectrograms.
[0727] "Means of conversion" refers to a process or device that converts acquired acoustic data into a format that is easily processed by a machine learning model.
[0728] A "multilayer learning model" is a neural network with multiple layers, designed to learn various intonation patterns for linguistic information.
[0729] "Linguistic information" refers to data about text spoken as audio and its context.
[0730] An "appropriate intonation pattern" is a pattern that shows appropriately adjusted pitch, volume, and rhythm to achieve natural and friendly conversation.
[0731] A "pre-trained model" is a deep learning model that has been trained in advance using data and is tuned to handle a specific task.
[0732] "Natural intonation" refers to intonation that is similar to the sounds humans actually make in real conversations, and is intended to provide a natural conversational experience.
[0733] "Audio data" refers to digital audio information that can be output as sound.
[0734] "Evaluation information" refers to data that shows the content of feedback and evaluations that users provide to the generated audio.
[0735] "Means of tuning a model" refers to the process of making modifications or updates to a machine learning model based on evaluation information in order to improve its performance.
[0736] This invention includes a process in which a server, a terminal, and a user cooperate to generate speech with natural intonation, as part of a voice dialogue system.
[0737] The server accesses a data storage device that stores acoustic feature data and retrieves the necessary data. This data is converted into formats such as MFCC (Mel-Frequency Cepstral Coefficients) and spectrograms, which serve to extract important features from the speech. The conversion is performed using dedicated speech processing software, and the data is preprocessed into a format suitable for machine learning. LibROSA, a Python-based library, is often used for this purpose.
[0738] Subsequently, the server uses deep learning frameworks such as TensorFlow and PyTorch to build a multi-layer learning model. This model is trained using pre-collected acoustic data with the goal of learning appropriate intonation patterns. Natural language processing techniques are also used in conjunction with intonation learning to take into account the meaning and sentiment of the text.
[0739] The user inputs text into the system via a terminal device. The input text is processed by the terminal and sent to the server as text data. An example of text is the prompt "Tell me the weather tomorrow." This text is converted into speech data with natural intonation by a trained deep learning model. As a result, the user can receive a natural and immersive response that closely resembles human dialogue.
[0740] The generated audio data is sent back to the device and played back to the user through speakers or headphones. This playback allows the user to evaluate the intonation and naturalness of the generated audio and provide feedback. This feedback is entered through the user interface and sent to the server. This information is used to adjust and improve the model, contributing to improved quality in subsequent audio generation processes.
[0741] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0742] Step 1:
[0743] The server retrieves acoustic feature data from the data storage device. In this step, audio recording data is retrieved from the database and output as an audio file format (e.g., WAV).
[0744] Step 2:
[0745] The server preprocesses the acquired acoustic data. This process uses an audio processing library such as LibROSA to convert the audio recording data into MFCCs or spectrograms. As a result, it outputs acoustic features as numerical data usable in machine learning models.
[0746] Step 3:
[0747] The server trains a multi-layer learning model using pre-processed data. At this stage, training is performed using TensorFlow or PyTorch, and the output model learns appropriate intonation patterns that take into account pitch and rhythm.
[0748] Step 4:
[0749] The user enters text into the system via the terminal. The entered text is presented as an example prompt, "Tell me tomorrow's weather," and sent to the server as text data.
[0750] Step 5:
[0751] The server inputs the received text data into a trained model and generates audio data with natural intonation. Specifically, it analyzes the text and outputs audio files (e.g., WAV) that reflect the emotion and context.
[0752] Step 6:
[0753] The device receives the generated audio data and plays it for the user via its playback function. In this step, the audio is output and played back through the speaker or headphones.
[0754] Step 7:
[0755] Users evaluate the naturalness and intonation of the played audio and provide feedback. This feedback is entered as an evaluation statement and sent to the server as data for improvement.
[0756] Step 8:
[0757] The server adjusts the model based on the feedback it receives. This feedback is used in the next speech generation process, where data calculations are performed to improve the model's accuracy and produce a more effective model.
[0758] (Application Example 1)
[0759] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0760] Conventional voice dialogue systems suffer from problems such as mechanical and unnatural intonation, making effective communication with users difficult. Furthermore, while there is a demand for natural voice dialogue in home electronic devices, current technology has its limitations. Therefore, there is a need for a means to achieve more natural and human-like voice dialogue.
[0761] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0762] In this invention, the server includes means for acquiring voice feature information from an information storage device and preprocessing this information, means for learning the optimal speech pattern for a sentence using a deep learning model, and means for generating voice information with natural speech based on the sentence input by the user using the trained model. This makes it possible for home electronic devices to provide more natural and effective conversational functions.
[0763] An "information storage device" is a device used to store data such as audio and text, and is used for data retrieval and preprocessing as needed.
[0764] "Voice feature information" refers to the basic attributes and patterns extracted from speech data, and in particular, includes features that contribute to natural intonation.
[0765] "Preprocessing" refers to the process of converting acquired data into a format suitable for analysis and learning, and includes data normalization and feature extraction.
[0766] A "deep learning model" is a model that uses a neural network consisting of multiple layers to learn complex patterns in data.
[0767] The "optimal vocalization pattern" refers to the tone and intonation of the voice that is suitable for achieving a natural and human-like intonation in response to the input text.
[0768] A "user" refers to a person who interacts with this system and utilizes its functions using voice input or text input.
[0769] "Text" refers to the text data to be processed, including the content input for speech generation.
[0770] A "trained model" refers to a model that has been trained using deep learning to specialize in specific patterns or tasks.
[0771] "Voice information with added pronunciation" refers to audio data generated based on text, using natural intonation.
[0772] A "household machine" is an automated machine used within the home, designed to provide a variety of functions, including conversational capabilities.
[0773] "Dialogue functionality" refers to a function that enables humans and machines to exchange information and communicate through voice and text.
[0774] This invention relates to a system for generating natural intonation of speech data. This system is based on three main components: a server, a terminal, and a user.
[0775] The server retrieves voice feature information from its data storage device. This includes previously collected audio data, which is preprocessed to make it suitable for training deep learning models. Specifically, it converts audio waveforms into acoustic features such as MFCCs and spectrograms. This conversion process utilizes servers equipped with NVIDIA GPU devices and TensorFlow software.
[0776] Next, the server constructs a deep learning model, such as a recurrent neural network (RNN) or its variant, the LSTM (Long Short-Term Memory) model, and trains it using pre-processed data. The trained model has the ability to add natural intonation to text and generates speech based on the text data entered by the user.
[0777] Users access the system through a device that allows for voice and text-based interaction. This device can be a smartphone, tablet, or even a robot as a household appliance. Text data provided by the user is sent to the server, and the natural-sounding voice generated by the server is then provided back to the user through the device.
[0778] A concrete example would be a scenario where a household appliance gently informs the user, "Today's weather is sunny," while they are getting ready in the morning, and then prompts them to confirm their daily schedule. An example of a prompt corresponding to this specific example would be, "Generate a response to the user's question with natural intonation."
[0779] This allows users to experience more human-like communication through home electronic devices.
[0780] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0781] Step 1:
[0782] The server retrieves voice feature information from the information storage device. This information includes past audio data and is in a format that allows mapping to text. The data is obtained from the original audio file and used as input for the next preprocessing step.
[0783] Step 2:
[0784] The server preprocesses the acquired audio data. This preprocessing includes conversion to MFCCs and spectrograms. This allows for the extraction of temporal and frequency features of the audio, resulting in high-dimensional audio features. The output is provided as a vector of acoustic features.
[0785] Step 3:
[0786] The server uses pre-processed speech features as input to build and train a deep learning model. Specifically, it uses a recurrent neural network (RNN) or LSTM model to learn the relationship between speech and text. In this step, the model's weights are updated to acquire the optimal speech pattern. Once training is complete, the model is ready to generate speech.
[0787] Step 4:
[0788] The user inputs text through their device and sends that data to the server. The input text data is provided as a sentence to which intonation should be added.
[0789] Step 5:
[0790] The server receives text sent by the user and generates speech using a trained model. The generation process converts the input text into speech data with optimal intonation. The output speech is then sent to the terminal as a digital audio file.
[0791] Step 6:
[0792] The device receives audio data sent from the server and plays it back to the user. This allows the user to hear audio with natural intonation.
[0793] Step 7:
[0794] Users provide feedback on the generated audio. This feedback concerns the naturalness and intelligibility of the audio and is sent to the server and recorded in a database.
[0795] Step 8:
[0796] The server readjusts the model based on user feedback. It analyzes the feedback and uses cross-validation to improve the model's accuracy. This step improves the overall system quality.
[0797] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0798] This invention provides a system for achieving natural intonation and user-centric dialogue in a voice dialogue system. This system consists of four elements: a server, a terminal, a user, and an emotion engine.
[0799] First, the server retrieves speech feature data from the database and preprocesses it. The speech data is processed into acoustic features and converted into a format suitable for training with a deep learning model. Intonation patterns are then learned based on this.
[0800] Next, during the training process using a deep learning model, the server learns the optimal intonation pattern for the user's input text. An emotion engine is incorporated, which can recognize emotions from the user's voice input and text. This engine uses emotion analysis technology to analyze keywords in the text and speech parameters to identify the user's emotional state.
[0801] When a user interacts using a device, the device sends data to the server in response to the user's input (text or voice). The data received by the server is then generated as speech data with natural intonation using a trained model. This generation process also takes into account the user's emotional state, and speech adjustments are made to suit that emotion.
[0802] The generated audio is sent to the device and played back to the user. At this time, the intonation adjusted by the emotion engine provides the user with a more natural and personalized conversational experience. Through this audio playback, the emotional nuances of the speech are more easily conveyed to the user.
[0803] Furthermore, user feedback and evaluations of the voice are sent from the device to the server. This feedback is shared with the emotion engine and used to improve the accuracy of the model and the engine. For example, if a user inputs "I'm feeling a little tired today," the emotion engine recognizes "fatigue," and the server generates a response with an intonation that takes this into consideration. This allows the user to experience more friendly and comfortable communication.
[0804] Thus, the present invention is a system that uses an emotion engine to realize natural voice dialogue optimized for the user.
[0805] The following describes the processing flow.
[0806] Step 1:
[0807] The server retrieves past audio data and its associated emotion labels from the database. This data is stored as text, corresponding audio files, and emotion information.
[0808] Step 2:
[0809] The server extracts acoustic features from the audio data and trains a deep learning model with emotion labels. The model learns the relationship between intonation patterns and emotional states in relation to speech, creating a foundation for natural speech generation.
[0810] Step 3:
[0811] The user inputs the text they want to speak through the device. Sometimes, the user's voice is also input at the same time.
[0812] Step 4:
[0813] The terminal sends user input text and voice to the server. An emotion engine also operates at this point to analyze the emotion in the voice.
[0814] Step 5:
[0815] The server uses an emotion engine to recognize emotions from the user's voice and adjusts the intonation generated by the deep learning model based on that emotion information.
[0816] Step 6:
[0817] The server generates audio data with the optimal intonation for the user's text. The generated audio takes into account the user's emotional state.
[0818] Step 7:
[0819] The server sends the generated voice data to the terminal. The voice, adjusted with emotional information, is more natural and approachable.
[0820] Step 8:
[0821] The device plays the generated audio to the user, who then receives feedback from the system.
[0822] Step 9:
[0823] Users provide feedback on the played audio, evaluating whether it is appropriate and what needs improvement.
[0824] Step 10:
[0825] The device sends user feedback to the server. This feedback is used for the continuous improvement of the emotion engine and deep learning models.
[0826] (Example 2)
[0827] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0828] In voice dialogue systems, generating voice responses with natural intonation that takes into account the user's emotions has been difficult with conventional technologies. Furthermore, there is a need for a system that effectively utilizes user feedback to improve model accuracy.
[0829] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0830] In this invention, the server includes means for preprocessing voice information obtained from a database, means for learning the optimal intonation pattern for the information using a machine learning algorithm, means for generating voice information with natural intonation using a trained model based on information input by the user, means for identifying the emotional state from the input information using sentiment analysis technology, and means for modifying the model based on the evaluation and emotional state identification results. This enables the generation of natural voice responses that correspond to the user's emotional state and continuous model improvement.
[0831] A "database" is a storage or management system that systematically organizes information to facilitate searching and access.
[0832] "Vocal information" refers to data related to human speech, and includes acoustic features, intonation, and parameters indicating emotion.
[0833] "Preprocessing" is the stage of shaping or transforming raw data into a format suitable for subsequent processing, and involves denoising data and extracting features.
[0834] A "machine learning algorithm" is a set of logic and procedures that allows a computer to learn patterns from data and perform predictions and classifications.
[0835] An "intonation pattern" is a structure that indicates the placement of intonation and accent in speech, and is important for achieving natural speech.
[0836] A "trained model" is an algorithm that has undergone an initial training process and incorporates knowledge based on diverse data.
[0837] "Sentiment analysis technology" is a technology that automatically identifies and classifies a person's emotions expressed in text or audio.
[0838] "Evaluation" refers to the act of users reviewing the results generated by a system and providing feedback based on quality and satisfaction.
[0839] "Modification" refers to making adjustments or changes to a system or algorithm in order to improve its initial state or its state after changes.
[0840] As an embodiment of this invention, details of a dialogue system using voice information are provided. The system consists of a server, terminals, users, and a data communication network connecting them.
[0841] The server acts as a central processing unit, handling audio information processing. Specifically, it retrieves audio information from the database and performs preprocessing. This preprocessing involves using audio processing libraries (e.g., Librosa) to remove noise and extract acoustic features. These features are then processed by deep learning models (e.g., RNNs and Transformers) to learn intonation patterns. The server also incorporates sentiment analysis technology, which uses natural language processing tools (e.g., NLTK and TextBlob) to identify emotions.
[0842] The user inputs voice or text via a terminal acting as an interactive device. The user's input is transmitted to the server through the terminal. The terminal then forwards the received voice data or text information directly to the server.
[0843] The device receives voice information generated from the server. This voice information is generated with intonation that takes the user's emotions into consideration, providing a more natural conversational experience. The user can provide feedback on this played voice information through the device.
[0844] The server accumulates user evaluations and feedback, and continuously improves the generated AI model. Cross-validation technology can be used to improve the model's accuracy.
[0845] As a concrete example, here is an example of a prompt: "Based on the text received from the user, generate a voice response with natural intonation that takes into account the emotional state." This prompt enables the system to engage in highly engaging conversations with the user, providing efficient and satisfying communication.
[0846] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0847] Step 1:
[0848] The server retrieves audio information from the database. It receives audio samples and associated metadata as input. This audio information contains various intonation patterns and is the target of analysis. The server uses an audio processing library to extract fundamental frequencies and spectral information and convert them into acoustic features. The output provides acoustic features for learning intonation.
[0849] Step 2:
[0850] The server inputs the acquired acoustic features into a deep learning model to learn the optimal intonation pattern. Machine learning algorithms are used here. During this process, the model acquires the ability to recognize patterns in information and identify intonation patterns. The learned intonation pattern is provided as output.
[0851] Step 3:
[0852] The user inputs data via voice or text through the device. This input includes data that reflects the user's intentions and emotions. The device sends this data to the server in preparation for the next processing step.
[0853] Step 4:
[0854] The server uses sentiment analysis technology to identify the user's emotional state based on input data received from the user. Input can be text or audio, along with a currently learned model. Here, natural language processing tools are used to analyze keywords and phrases to determine the user's emotions. The identified emotional state is provided as output.
[0855] Step 5:
[0856] The server uses a trained model to adjust intonation according to the identified emotional state and generate a voice response. The input consists of trained intonation patterns and the user's emotion. The output is voice data with a natural intonation that matches the user's emotion.
[0857] Step 6:
[0858] The device receives the generated audio and plays it for the user. It handles audio data sent from the server as input. During playback, it ensures that the intonation is natural and aligns with the user's emotions. The final audio experience delivered to the user is then produced as output.
[0859] Step 7:
[0860] Users provide feedback on voice responses. As input, they enter their evaluation or comments on the generated voice into the device. The device sends this to the server, which helps improve the model. As output, the feedback is accumulated and used to generate the next response.
[0861] (Application Example 2)
[0862] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0863] Conventional voice dialogue systems have struggled to provide natural intonation that fully considers the user's emotions, making it difficult to achieve emotionally resonant dialogue. This problem makes it challenging to provide friendly and comfortable communication for users.
[0864] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0865] In this invention, the server includes means for acquiring speech feature information from a data set and preprocessing it, means for learning the optimal intonation pattern for the recording using a machine learning system, and means for analyzing the user's emotions and reflecting emotion-based intonation in the speech information. This makes it possible to provide emotion-adaptive speech dialogue with natural intonation.
[0866] A "data set" is a collection of data that is gathered and managed for a specific purpose.
[0867] "Audio feature information" refers to information containing features such as frequency, intensity, and pitch extracted from audio data.
[0868] "Preprocessing" is the process of transforming and processing data to make it suitable for analysis and learning.
[0869] A "machine learning system" is a system that uses algorithms to learn rules and patterns from data and perform predictions and classifications.
[0870] "Records" are data or documents that are stored and protected for later reference.
[0871] An "intonation pattern" is a pattern that shows the combination and arrangement of intonation in the sounds of a language.
[0872] A "user" is a person who uses a system or device.
[0873] "Emotional analysis" is the process of extracting emotional states from audio or text as numerical values or labels.
[0874] "Audio information" refers to all information transmitted as sound, including audio signal data.
[0875] "Emotionally adaptive voice interaction" refers to voice interaction that recognizes the user's emotions and engages in natural conversation tailored to those emotions.
[0876] The system implementing this invention consists of three elements: a server, a terminal, and a user. The server acquires speech feature information from a data set and preprocesses that data. The preprocessed data is used by a machine learning system to learn the optimal intonation pattern for recording. In particular, TensorFlow / Keras is used as a deep learning framework to learn natural intonation. Speech processing tools such as PyDub are utilized for processing the speech data.
[0877] The server then generates speech information with natural intonation using a trained system based on the user's input. The terminal receives the user's input and sends that information to the server. The server analyzes the user's emotions using an emotion recognition engine. Using the natural language processing library NLTK, it determines the user's emotions and generates speech information with the appropriate intonation.
[0878] The generated voice information is sent to the terminal and played back by the user. Through the played voice, emotionally adaptive voice dialogue is realized. This dialogue can be used particularly in home assist robots and improves the quality of everyday communication.
[0879] As a specific example, in a morning scenario, when the robot says, "Good morning. How are you feeling today?", and the user inputs, "I'm feeling a little down," the robot will provide an emotionally adaptive response. In this case, the robot will reply in a gentle tone of voice, "I see. Let's do something to help you relax."
[0880] An example of a prompt is, "Generate a kind and considerate intonation pattern for when a user says, 'Today was tough.'" This sentence is used in a generative AI model to generate responses appropriate to the user's emotions.
[0881] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0882] Step 1:
[0883] The terminal acquires text and voice data entered by the user. This data is sent to the server as input data. The input data is expected to contain the user's intentions and emotions.
[0884] Step 2:
[0885] The server extracts speech feature information from the received speech data and performs natural language processing on the text data as needed. It extracts acoustic features from the speech data and performs emotion recognition on the text data. Based on the input data, it analyzes the features of the speech data and generates a feature vector as an intermediate output.
[0886] Step 3:
[0887] The server uses a deep learning model to generate the optimal intonation pattern from the feature vector. Here, the learned pattern is applied using a generative AI model. The output of this step is the generated intonation pattern.
[0888] Step 4:
[0889] The server uses an emotion recognition engine to analyze the user's emotions and adjust the intonation accordingly. Specifically, it takes the emotion analysis results as input and adjusts the intonation appropriately. It generates speech data with intonation that matches the emotion.
[0890] Step 5:
[0891] The server sends the generated audio data to the terminal. The terminal plays this data back to the user. This provides the user with natural and emotionally appropriate voice responses. The final output is the intoned audio played back to the user.
[0892] Step 6:
[0893] The user evaluates the played audio and inputs their feedback into the device. The device then sends the feedback to the server. This evaluation information is used to further improve the system.
[0894] Step 7:
[0895] Based on the feedback received, the server adjusts the speech generation system and improves the accuracy of the learning model and emotion engine in preparation for the next use. This improves the overall performance of the system.
[0896] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0897] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0898] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0899] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0900] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0901] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0902] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0903] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0904] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0905] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0906] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0907] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0908] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0909] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0910] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0911] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0912] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0913] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0914] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0915] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0916] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0917] The following is further disclosed regarding the embodiments described above.
[0918] (Claim 1)
[0919] A means for obtaining speech feature data from a database and preprocessing it,
[0920] A method for learning the optimal intonation pattern for text using a deep learning model,
[0921] A means for generating speech data with natural intonation using a trained model based on text input by a user,
[0922] A means of playing the generated audio for the user and receiving their evaluation,
[0923] A system that includes means for modifying the model based on evaluations.
[0924] (Claim 2)
[0925] The system according to claim 1, comprising means for performing cross-validation of data to improve the accuracy of the model.
[0926] (Claim 3)
[0927] The system according to claim 1, comprising means for recording feedback information for generated audio data and utilizing the feedback in the next generation process.
[0928] "Example 1"
[0929] (Claim 1)
[0930] A means for acquiring acoustic feature data from a data storage device and performing a transformation,
[0931] A method for learning appropriate intonation patterns for linguistic information using a multi-layer learning model,
[0932] A means for generating acoustic data with natural intonation using a trained model based on language information input by the user,
[0933] A means of playing back the generated sound to the user and receiving evaluation information,
[0934] A system that includes means for adjusting a model based on evaluation information.
[0935] (Claim 2)
[0936] The system according to claim 1, comprising means for performing cross-validation of data and improving the accuracy of the model.
[0937] (Claim 3)
[0938] The system according to claim 1, comprising means for recording evaluation information for generated acoustic data and using the evaluation information in the next generation process.
[0939] "Application Example 1"
[0940] (Claim 1)
[0941] A means for acquiring voice feature information from an information storage device and preprocessing said information,
[0942] A method for learning the optimal pronunciation pattern for a given text using a deep learning model,
[0943] A means for generating voice information with natural-sounding pronunciation based on text input by a user, using a pre-trained model,
[0944] A means of playing back the generated voice to the user and receiving their evaluation,
[0945] A means of modifying the model based on the evaluation,
[0946] A system that includes means for providing a user with an interactive function to assist them in their daily activities.
[0947] (Claim 2)
[0948] The system according to claim 1, comprising means for performing cross-validation of data and improving the accuracy of the model.
[0949] (Claim 3)
[0950] The system according to claim 1, comprising means for recording response information to generated voice information and using the response in the next generation operation.
[0951] "Example 2 of combining an emotion engine"
[0952] (Claim 1)
[0953] A means for preprocessing audio information obtained from a database,
[0954] A method for learning the optimal intonation pattern for information using machine learning algorithms,
[0955] A means for generating speech information with natural intonation using a trained model based on information input by the user,
[0956] A means of playing the generated audio for the user and receiving their evaluation,
[0957] A means of identifying emotional states from input information using emotion analysis technology,
[0958] A system that includes means for modifying a model based on the results of evaluation and identification of emotional states.
[0959] (Claim 2)
[0960] The system according to claim 1, comprising means for performing cross-validation of data to improve the accuracy of the model.
[0961] (Claim 3)
[0962] The system according to claim 1, comprising means for recording opinion information regarding generated audio information and using the opinion in the next generation process.
[0963] "Application example 2 when combining with an emotional engine"
[0964] (Claim 1)
[0965] A means for obtaining speech feature information from a data set and preprocessing it,
[0966] A means of learning the optimal intonation pattern for a recording using a machine learning system,
[0967] A means for generating speech information with natural intonation using a trained system based on records entered by the user,
[0968] A means of playing the generated audio for the user and receiving their evaluation,
[0969] A means of modifying the system based on the evaluation,
[0970] A means of analyzing the user's emotions and reflecting emotion-based intonation in the audio information,
[0971] A means of providing communication support by generating emotionally adaptive responses to users through voice dialogue,
[0972] A system that includes this.
[0973] (Claim 2)
[0974] The system according to claim 1, comprising means for performing data cross-validation to improve the accuracy of the system.
[0975] (Claim 3)
[0976] The system according to claim 1, comprising means for recording feedback information for generated audio information and utilizing the feedback in the next generation process. [Explanation of Symbols]
[0977] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring voice feature information from an information storage device and preprocessing said information, A method for learning the optimal pronunciation pattern for a given text using a deep learning model, A means for generating voice information with natural-sounding pronunciation based on text input by a user, using a pre-trained model, A means of playing back the generated voice to the user and receiving their evaluation, A means of modifying the model based on the evaluation, A system that includes means for providing a user with an interactive function to assist them in their daily activities.
2. The system according to claim 1, comprising means for performing cross-validation of data and improving the accuracy of the model.
3. The system according to claim 1, comprising means for recording response information to generated voice information and using the response in the next generation operation.