system

The system addresses the limitations of existing voice reproduction systems by synchronizing voice and mouth movements, providing natural and secure conversational experiences with distant individuals.

JP2026062184APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing voice reproduction systems lack naturalness and realism, failing to provide a satisfying user experience due to limitations in voice and mouth movement synchronization, especially for individuals who rarely communicate with distant family or friends, and there is a need for systems that offer psychological security and familiarity.

Method used

A system that collects voice samples, analyzes them to extract features, trains a deep learning model, generates responses, and synchronizes mouth movements with the avatar's display, allowing natural conversational experiences.

Benefits of technology

Enables natural and psychologically secure conversations with distant individuals by accurately reproducing voice and mouth movements, enhancing user familiarity and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062184000001_ABST
    Figure 2026062184000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for collecting voice samples from a specific person in a remote location, A means for analyzing collected audio samples and extracting audio features, A method for training a deep learning model using speech features, A means for generating a response in response to voice input from a user using a trained deep learning model, A means for outputting the generated response as audio, A means of displaying the avatar's mouth movements in sync with the outputted audio, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a need for a technology that can reproduce the voice and speaking style of a specific person for people who feel it difficult to frequently communicate with family members and close friends who are in a distant place. Also, there is a need for a system that can provide a user with a psychological sense of security and familiarity by offering a conversation experience with a deceased person. However, the conventional voice reproduction systems have lacked in the naturalness and realism of voices, and thus have not been able to obtain sufficient satisfaction. Furthermore, the technology for reproducing mouth movements synchronized with voices has also been limited, making it difficult to improve the user experience.

Means for Solving the Problems

[0005] The present invention provides a system that includes means for collecting voice samples of a specific person in a remote location, means for analyzing the collected voice samples to extract voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from a user using the trained deep learning model, means for outputting the generated response as voice, and means for displaying the mouth movements of an avatar in synchronization with the output voice. This makes it possible to reproduce the pronunciation patterns and intonation of a specific person, and furthermore, to naturally express the mouth movements of the avatar in accordance with the voice. As a result, it is possible to provide a natural conversational experience with family, close friends, or deceased loved ones in remote locations, and to make the user feel psychologically secure and familiar with them.

[0006] A "voice sample" is a recording of audio data spoken by a specific person.

[0007] "Speech features" are characteristic elements extracted from speech data, such as pronunciation patterns, intonation, pitch, speed, and intonation.

[0008] A "deep learning model" is an algorithm based on a neural network that has a multi-layered learning structure and can learn complex patterns from large amounts of data.

[0009] "Training" is the process of using data to teach a model specific patterns or features.

[0010] "Response" refers to the content of the conversation or reply that the system generates in response to user input.

[0011] Text-to-speech (TTS) is a technology that converts text data into speech data.

[0012] "Lip sync" is a technology that naturally reproduces the mouth movements of an avatar in sync with audio data.

[0013] An "avatar" is a video of a person or character created using computer graphics.

[0014] Morphological analysis is a technique that breaks down text into words and morphemes and analyzes their meaning and function.

[0015] "Analyzing intent" is the process of understanding a user's intentions and objectives from the content of their spoken words or text. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The present invention mainly involves the following processing steps.

[0038] 1. Collection of audio samples:

[0039] Voice samples are collected when users speak to a smart speaker.

[0040] The device (smart speaker) records the user's voice and sends that data to the server.

[0041] For example, when a user says "Good morning," the voice is recorded by the smart speaker, and the voice data is sent to the server.

[0042] 2. Storage and analysis of audio data:

[0043] The server temporarily stores the received audio data.

[0044] The saved audio data is converted into text using speech recognition technology. Additionally, audio features (pronunciation patterns, intonation, etc.) are extracted.

[0045] The extracted speech features are stored in a database and used later for deep learning.

[0046] 3. Training deep learning models:

[0047] The server uses the accumulated audio features to train a deep learning model.

[0048] During the training process, specific speaking styles and intonations are learned, and models are built for generating natural dialogue.

[0049] For example, phrases that users repeatedly say, such as "Hello" or "How are you?", can be used as training data.

[0050] 4. Receiving and analyzing questions and conversations:

[0051] When a user speaks a question or engages in a conversation with a smart speaker, the device records it and sends it to a server.

[0052] The server analyzes the received audio, converts it to text, and performs morphological analysis to understand the user's intent.

[0053] For example, if a user asks "What did you do today?", the content of that question is converted into text and analyzed.

[0054] 5. Response generation and speech synthesis:

[0055] A trained deep learning model is used to generate appropriate responses to user questions.

[0056] The generated response is converted into speech data using text-to-speech (TTS) technology.

[0057] For example, a response like "I had a relaxing day today" might be generated.

[0058] 6. Lip sync and avatar display:

[0059] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0060] The terminal receives voice data and lip-sync data, and uses them to actually respond to the user.

[0061] The user can hear the response and simultaneously see the avatar's natural mouth movements.

[0062] For example, an avatar saying "I had a relaxing day today" will be displayed with accurate mouth movements.

[0063] Thus, the present invention provides a system that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away, by combining a technology that reproduces the voice characteristics of a specific person with a lip-sync technology that synchronizes with the voice. It also includes specific means to make the user feel psychologically secure and familiar with the system.

[0064] The following describes the processing flow.

[0065] Step 1:

[0066] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[0067] Step 2:

[0068] The device (smart speaker) records the user's voice and saves it to a temporary buffer. After recording is complete, it sends the voice data to the server.

[0069] Step 3:

[0070] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[0071] Step 4:

[0072] The server converts audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features, including pronunciation patterns and intonation.

[0073] Step 5:

[0074] The server stores the extracted audio features in a database. Simultaneously, it uses this data to train a deep learning model. A multi-layered neural network is used for training.

[0075] Step 6:

[0076] The user initiates a new conversation or asks a question to the smart speaker. For example, they might ask, "How was your day?"

[0077] Step 7:

[0078] The device (smart speaker) re-records the user's voice and sends it to the server.

[0079] Step 8:

[0080] The server receives the audio data and converts it into text data. It then performs morphological analysis on the text to analyze the user's intent.

[0081] Step 9:

[0082] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a response like, "I had a relaxing day."

[0083] Step 10:

[0084] The server converts the generated text response into speech data using text-to-speech (TTS) technology. This enables natural-sounding speech.

[0085] Step 11:

[0086] The server generates lip-sync data corresponding to the audio data. This allows the audio and the avatar's mouth movements to be synchronized.

[0087] Step 12:

[0088] The server sends audio data and lip-sync data to the terminal.

[0089] Step 13:

[0090] The device (smart speaker) receives audio data and lip-sync data from the server. Based on the received data, it plays the audio and simultaneously displays the mouth movements of the avatar.

[0091] Step 14:

[0092] Users listen to responses generated through smart speakers and observe the natural mouth movements of their avatars. This allows them to enjoy a natural conversational experience with loved ones.

[0093] In this way, a system is created that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away.

[0094] (Example 1)

[0095] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] Conventional speech recognition and dialogue generation systems struggle to reproduce the natural speaking style and intonation of a specific person located remotely. As a result, they are insufficient as systems that provide users with a sense of psychological security and familiarity. Furthermore, these systems often lack visual feedback for the generated speech, leading to mismatches between speech and mouth movements. This, in turn, negatively impacts the user experience.

[0097] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0098] In this invention, the server includes means for acquiring voice data, means for storing the acquired voice data, means for converting the stored voice data into text using speech recognition technology, means for extracting and accumulating voice features, means for training a deep learning model using the accumulated voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice using speech synthesis technology, and means for displaying the mouth movements of an avatar in synchronization with the output voice. This makes it possible to generate a response in which the voice and mouth movements are naturally synchronized, providing the user with a sense of psychological security and familiarity.

[0099] "Means of acquisition" refers to devices or mechanisms for collecting the user's voice.

[0100] "Means of storage" refers to a storage device or service for temporarily storing the collected audio data.

[0101] "Speech recognition technology" refers to algorithms and software used to convert speech data into text data.

[0102] "Speech features" are identifiable characteristics such as pronunciation patterns and intonation that are extracted from speech data.

[0103] "Means of storage" refer to databases or storage devices for saving and managing extracted speech features.

[0104] A "deep learning model" is a neural network model that is trained using speech features.

[0105] "Means for generating responses in response to voice input" refers to a mechanism that uses a trained deep learning model to create appropriate responses to user voice input.

[0106] "Speech synthesis technology" refers to algorithms and software used to convert generated text data into speech data.

[0107] "Means for displaying the mouth movements of an avatar" refers to a display device or software that reproduces the mouth movements of an avatar in synchronization with the outputted audio.

[0108] To implement this invention, three elements are necessary: ​​a user, a terminal (smart speaker), and a server. The user speaks to the terminal, the terminal sends voice data to the server, and the server analyzes and processes the voice data to generate an appropriate response.

[0109] When a user speaks into a smart speaker, the device uses its microphone to capture the user's voice. Specifically, the user might say things like "Good morning" or "What did you do today?". This voice data is captured by the device and sent to a server via the internet.

[0110] The transmitted audio data is first temporarily stored on the server. The server converts the stored audio data into text using speech recognition technology. Services such as Google's Speech-to-Text API are used for this process. The text data obtained through speech recognition technology is stored in a database on the server.

[0111] Next, the server extracts speech features (e.g., pronunciation patterns, intonation) from the audio data. These speech features are used as data to train a deep learning model (e.g., TENSORFLOW®). During the training process, the server learns the user's unique speaking style and intonation, and builds a model for generating natural dialogue.

[0112] When the user speaks a question or has a conversation into the device again, the device records the audio and sends it to the server. The server converts the audio back into text and performs morphological analysis (e.g., using MeCab). Based on the analyzed text data, the server uses a trained deep learning model to generate an appropriate response. For example, it might generate a response such as, "I had a relaxing day."

[0113] The generated text responses are converted into audio data using speech synthesis technology. Speech synthesis services such as the Google Text-to-Speech API are used for this process. The server then generates lip-sync data synchronized with the audio data to recreate the avatar's mouth movements.

[0114] The device uses the received audio and lip-sync data to respond to the user. The user can listen to the avatar's response, such as "I had a relaxing day," while observing the avatar's natural mouth movements.

[0115] Specific example

[0116] One day, a user asks a smart speaker, "What's the weather like tomorrow?" The device records the voice and sends it to a server. The server converts the voice to text, performs further analysis, and generates an appropriate response. This response is "It will be sunny tomorrow," and is converted into voice data via speech synthesis technology. The generated voice data and lip-sync data are sent to the device, and an avatar speaks "It will be sunny tomorrow," with its mouth movements naturally reproduced.

[0117] Example of a prompt

[0118] 1. "What's the weather like tomorrow?"

[0119] 2. "Please tell me about the latest news."

[0120] 3. "What fun things happened to you today?"

[0121] This system generates responses where voice and mouth movements are naturally synchronized, providing users with a sense of psychological reassurance and familiarity.

[0122] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0123] Step 1:

[0124] The user speaks to the smart speaker.

[0125] Input: Voice (Example: "Good morning")

[0126] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0127] Output: Recorded audio data

[0128] Step 2:

[0129] The device (smart speaker) sends the recorded audio data to the server.

[0130] Input: Recorded audio data

[0131] Specific operation: Captured audio data is sent to the server via the internet.

[0132] Output: Audio data sent to the server

[0133] Step 3:

[0134] The server temporarily stores the received audio data.

[0135] Input: Audio data sent to the server

[0136] Specific operation: The audio data is saved to storage (e.g., Amazon S3).

[0137] Output: Saved audio data

[0138] Step 4:

[0139] The server uses speech recognition technology to convert the audio data into text.

[0140] Input: Saved audio data

[0141] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and text data is returned.

[0142] Output: Text data

[0143] Step 5:

[0144] The server extracts speech features from the audio data and stores them in a database.

[0145] Input: Saved audio data

[0146] Specific operation: Audio data is analyzed, features such as pronunciation patterns and intonation are extracted, and these are stored in a database.

[0147] Output: Speech feature data

[0148] Step 6:

[0149] The server uses speech features to train a deep learning model.

[0150] Input: Speech feature data

[0151] Specific operation: A neural network is trained using a framework such as TensorFlow with audio feature data.

[0152] Output: Trained deep learning model

[0153] Step 7:

[0154] The user speaks to the smart speaker, asking questions or engaging in conversation.

[0155] Input: Voice (Example: "What did you do today?")

[0156] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0157] Output: Recorded audio data

[0158] Step 8:

[0159] The device sends the recorded audio data to the server.

[0160] Input: Recorded audio data

[0161] Specific operation: Captured audio data is sent to the server via the internet.

[0162] Output: Audio data sent to the server

[0163] Step 9:

[0164] The server converts the audio data into text and performs morphological analysis.

[0165] Input: Audio data sent to the server

[0166] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and after obtaining text data, morphological analysis is performed using MeCab, etc.

[0167] Output: Morphologically analyzed text data

[0168] Step 10:

[0169] The server uses a trained deep learning model to generate appropriate responses to questions and conversations.

[0170] Input: Morphologically analyzed text data

[0171] Specific operation: Input text data into a trained deep learning model and generate an appropriate response.

[0172] Output: Generated text response

[0173] Step 11:

[0174] The server converts the generated text response into speech data using speech synthesis technology.

[0175] Input: Generated text response

[0176] Specific operation: Convert text responses into speech data using the Google Text-to-Speech API, etc.

[0177] Output: Generated audio data

[0178] Step 12:

[0179] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0180] Input: Generated audio data

[0181] Specific operation: Analyzes audio data and creates lip-sync data to match the audio.

[0182] Output: Lip-sync data

[0183] Step 13:

[0184] The device receives voice data and lip-sync data and displays a response to the user.

[0185] Input: Generated audio data and lip-sync data

[0186] Specific operation: Based on the received data, the avatar responds to the user with voice and simultaneously displays the movement of its mouth.

[0187] Output: Voice responses and mouth movements by the avatar

[0188] In this way, through each processing step, users can enjoy natural conversations with specific individuals in remote locations through the reproduction of natural dialogue and the accompanying mouth movements of their avatars.

[0189] (Application Example 1)

[0190] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0191] Conventional speech recognition and voice response systems have limitations in achieving natural dialogue with users, and systems combining speech synthesis and lip-syncing have not been widely implemented. Furthermore, in food delivery services, it has been difficult to provide personalized menu suggestions and answer questions through dialogue with users. This invention aims to solve these problems and enable users to use food delivery services more comfortably.

[0192] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0193] In this invention, the server includes means for collecting voice samples from a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, virtual assistant means for analyzing the user's questions, generating an appropriate response, and outputting that response using speech synthesis technology, and means for using a generative AI model that generates a response based on the analysis results of the user's voice input. This enables a more natural and intuitive conversation when a user uses a food delivery service.

[0194] A "voice sample" is data that is a recording of a person speaking in a remote location.

[0195] "Speech features" refer to characteristic information such as pronunciation patterns and intonation extracted from collected speech samples.

[0196] A "deep learning model" is a model that uses a multi-layered neural network to learn from large amounts of data and perform predictions and classifications on new data.

[0197] A "virtual assistant" is a software system that generates appropriate responses to user voice input and outputs those responses using speech synthesis technology.

[0198] A "generative AI model" is an artificial intelligence model trained to generate responses based on a user's voice input.

[0199] A "lip-syncing method" is a means of displaying the mouth movements of an avatar in sync with the generated audio.

[0200] A "food delivery service" is a service that delivers food to customers based on their orders.

[0201] This invention is a system that collects voice samples from people in remote locations and uses them to provide users with natural, real-time conversations. This system is particularly effective in food delivery services, where users can use devices such as smartphones or smart glasses to interact with a virtual assistant and obtain information about menus and recommended dishes.

[0202] First, the server collects voice samples and temporarily stores the data. These voice samples are obtained when the user speaks into their smartphone or smart glasses. For example, if the user asks, "What pasta dish do you recommend?", the voice is recorded by the device and sent to the server.

[0203] Next, the server analyzes the collected audio data and converts it into text using speech recognition technology. Simultaneously, it extracts audio features (such as pronunciation patterns and intonation) and stores them in a database. This allows a deep learning model to be trained using the audio features.

[0204] A trained deep learning model generates appropriate responses to voice input from the user. When the user provides voice input again, the server analyzes the speech to understand the intent and uses the generative AI model to generate an appropriate response. This response is then converted into audio data using speech synthesis technology and output to the user.

[0205] Furthermore, lip-sync data synchronized with this generated audio data is created, allowing the virtual assistant avatar to reproduce natural mouth movements. This allows users to visually confirm that the virtual assistant is speaking naturally.

[0206] As a concrete example, when a user asks, "What pasta dish do you recommend?", the system generates a response through the following process: It replies with voice, "My recommendation is carbonara," and an avatar makes the same response with natural mouth movements. This allows the user to experience an immersive conversation.

[0207] The hardware used will be a smartphone, smart glasses, and a server. The software used will be speech recognition technology (e.g., the speech_recognition library), speech synthesis technology (e.g., the gtts library), and a generative AI model (e.g., the GPT-2 model from the transformers library).

[0208] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0209] Step 1:

[0210] When a user speaks to their smartphone or smart glasses, their voice is recorded by the device. The recorded voice data is then sent directly to the server. For example, if a user asks, "What pasta dish do you recommend?", that voice will be collected. The input is the user's voice data, and the output is the voice data sent to the server.

[0211] Step 2:

[0212] The server temporarily stores the received audio data. Then, it converts the audio data into text data using speech recognition technology. Here, the `speech_recognition` library is used to convert speech to text. The input is audio data, and the output is text data.

[0213] Step 3:

[0214] The server further analyzes the text data and extracts speech features (pronunciation patterns, intonation, etc.). A speech feature extraction algorithm is used for the analysis, and the extracted features are stored in a database. The input is text data, and the output is speech feature data.

[0215] Step 4:

[0216] The server accumulates speech feature data and uses it to train a deep learning model. A large amount of speech features are used for training, and a trained generative AI model is constructed. The input is speech feature data, and the output is the trained generative AI model.

[0217] Step 5:

[0218] When the user speaks a question to the device again, the device records the audio and sends it to the server. The server analyzes the received audio, converts it into text data, and performs morphological analysis to understand the user's intent. The input is the newly collected audio data, and the output is the analyzed text data and the results of the intent analysis.

[0219] Step 6:

[0220] The server uses a trained generative AI model to generate appropriate responses to user voice input. The generated responses are converted into speech data using speech synthesis technology (e.g., the GTTS library). The input is the intent analysis result, and the output is the speech response data.

[0221] Step 7:

[0222] The server generates lip-sync data synchronized with the generated audio data, ensuring that the virtual assistant avatar reflects natural mouth movements. The lip-sync data is used to control the avatar's mouth movements in real time. The input is audio data, and the output is lip-sync data.

[0223] Step 8:

[0224] The terminal receives audio data and lip-sync data from the server, plays an audio response to the user, and displays the avatar's lip-sync. The input is audio data and lip-sync data, and the output is the audio response to the user and the avatar's natural mouth movements.

[0225] Through the steps outlined above, users can ask questions about the food delivery service in a natural, conversational format. This allows users to obtain detailed information through dialogue, significantly improving convenience.

[0226] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0227] This invention is a system that reproduces the voice characteristics of a specific person in a remote location, recognizes the user's emotions, and adjusts the response accordingly. The system consists of a smart speaker and a server, and each process is performed as follows.

[0228] 1. Collection and storage of audio samples:

[0229] The user speaks to the smart speaker, which provides a voice sample of a specific person. For example, the user might say "Good morning."

[0230] The device (smart speaker) records the user's voice and sends that voice data to the server.

[0231] 2. Analysis of audio data:

[0232] The server stores the received audio data and converts it into text data using speech recognition technology. Based on the converted text, it extracts audio features such as pronunciation patterns and intonation.

[0233] The extracted speech features are stored in a database and used to train a deep learning model.

[0234] 3. Training deep learning models:

[0235] The server uses the accumulated speech features to train a deep learning model to learn specific speaking styles and intonations.

[0236] The trained model will be used to generate future user responses.

[0237] 4. User emotion recognition:

[0238] When a user speaks a new question or engages in a conversation with the smart speaker, the device records the audio again and sends it to the server.

[0239] The server analyzes the received audio data and uses emotion recognition technology to analyze the user's emotions.

[0240] 5. Analyzing the intent behind questions and conversations:

[0241] The server converts the audio data into text, performs morphological analysis, and understands the user's intent.

[0242] 6. Response generation and adjustment:

[0243] Using a pre-trained deep learning model, it generates appropriate responses to user questions. For example, it can generate a response like, "I had a relaxing day today."

[0244] The response is adjusted based on the user's emotions as assessed by the emotion recognition engine. For example, if the user appears sad, it might generate a response such as, "I had a relaxing day today. You seem a little down, are you okay?"

[0245] The generated text responses are converted into speech data using text-to-speech (TTS) technology.

[0246] 7. Lip sync and avatar display:

[0247] The server generates lip-sync data corresponding to the audio data, naturally expressing the avatar's mouth movements and facial expressions.

[0248] The device receives audio data and lip-sync data, plays the audio, and displays the avatar's facial expressions and mouth movements.

[0249] 8. User experience:

[0250] Users can listen to the responses generated through the smart speaker and see the avatar's natural mouth movements and facial expressions.

[0251] This allows users to enjoy natural conversations with family members or deceased loved ones who live far away, and to receive responses that are sensitive to the user's emotions.

[0252] In this way, the present invention realizes a system that realistically reproduces the voice characteristics of a specific person, understands the user's emotions, and provides a more natural and approachable conversational experience.

[0253] The following describes the processing flow.

[0254] Step 1:

[0255] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[0256] Step 2:

[0257] The device (smart speaker) records the user's voice and saves it to a temporary buffer. It then sends the recorded voice data to the server.

[0258] Step 3:

[0259] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[0260] Step 4:

[0261] The server converts the audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features (such as pronunciation patterns and intonation).

[0262] Step 5:

[0263] The server stores the extracted speech features in a database. Simultaneously, this data is used to train a deep learning model.

[0264] Step 6:

[0265] The user initiates a new question or conversation with the smart speaker. For example, they might ask, "How was your day?"

[0266] Step 7:

[0267] The device (smart speaker) re-records the user's voice and sends it to the server.

[0268] Step 8:

[0269] The server analyzes the received audio data and converts it back into text data. This text is then subjected to morphological analysis to determine the user's intent.

[0270] Step 9:

[0271] The server uses an emotion recognition engine during the process of analyzing voice data to analyze the user's emotions (such as joy, anger, sadness, etc.).

[0272] Step 10:

[0273] The server uses a trained deep learning model to evaluate the user's emotions and generate appropriate responses. For example, it generates a response like "I've been having a relaxing day."

[0274] Step 11:

[0275] Based on the emotion recognition result, the server adjusts the response content. For example, if the user looks sad, it changes the content to something like "I've been having a relaxing day. You seem a bit down. Are you okay?"

[0276] Step 12:

[0277] The generated text response is converted into audio data using text-to-speech (TTS) technology. This makes the response a natural speech.

[0278] Step 13:

[0279] The server generates lip sync data corresponding to the audio data and creates data for naturally expressing the movements and expressions of the avatar's mouth.

[0280] Step 14:

[0281] The server sends the audio data and lip sync data to the terminal.

[0282] Step 15:

[0283] The terminal (smart speaker) receives the audio data and lip sync data from the server. It plays the audio and displays the expressions and mouth movements of the avatar.

[0284] Step 16:

[0285] The user listens to the generated response and checks the natural mouth movements and expressions of the avatar. This allows the user to enjoy a natural conversation experience with family members who are far away or with deceased individuals.

[0286] With this specific processing flow, the present invention faithfully reproduces the voice characteristics and pronunciation patterns of a specific person and provides a friendly response corresponding to the user's emotions.

[0287] (Example 2)

[0288] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".

[0289] It is difficult to provide a natural and friendly conversation experience with a specific person located remotely. Also, it is not easy to recognize the user's emotions and generate an appropriate response accordingly. Therefore, it is an issue to accurately reproduce the voice and speaking style of a specific person while making a response that conforms to the user's emotions.

[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0291] In this invention, the server includes means for collecting voice samples of a specific person located remotely, means for analyzing the collected voice samples to extract voice feature quantities such as pronunciation patterns and intonation, means for training a deep learning model using the voice feature quantities, means for recognizing the user's emotions and adjusting the response content, and means for converting the voice input from the user into text and analyzing the intention. Thereby, it becomes possible to realistically reproduce the voice of a specific person and provide a natural conversation experience corresponding to the user's emotions.

[0292] A "voice sample" is a digital recording of the voice uttered by the user and is an imitation of the voice of a specific person.

[0293] A "pronunciation pattern" refers to features such as the flow of phonemes, rhythm, stress, and accentuation in speech.

[0294] "Intonation" refers to the changes in pitch and cadence of the voice in spoken language.

[0295] "Speech features" are specific data points extracted from speech samples, and include elements such as pronunciation patterns, intonation, pitch, and length.

[0296] A "deep learning model" is a type of artificial intelligence trained using large amounts of data, and specifically refers to models that use neural networks.

[0297] "Emotion recognition" refers to technology that analyzes and identifies a user's emotional state from their voice or text.

[0298] An "avatar" refers to a representation of a person or character displayed in a virtual space for interaction with a user.

[0299] "Lip sync" refers to a technology that synchronizes the mouth movements of an avatar with the audio.

[0300] "Speech recognition technology" is a technology for converting speech into text, and includes the process of analyzing a speech sample and converting it into text data.

[0301] "Speech synthesis technology" refers to the technology that generates speech based on text data.

[0302] Modes for carrying out the invention

[0303] This invention relates to a system that reproduces the voice characteristics of a specific person located remotely, recognizes the user's emotions, and adjusts its response accordingly. The system consists of a smart speaker and a server and includes the following elements:

[0304] 1. Collection and storage of audio samples

[0305] The user collects voice samples by speaking to the smart speaker. For example, the user utters "Good morning". The terminal (smart speaker) records the user's voice with a microphone and temporarily stores this voice data in the internal memory. Next, the terminal transmits the voice data to the server via Wi-Fi.

[0306] 2. Analysis of Voice Data

[0307] The server stores the received voice data in a dedicated database. Then, it uses voice recognition technology to convert the voice data into text data. Specifically, it utilizes the Google Speech-to-Text API to convert it into text data. Based on this text data, it extracts voice feature quantities such as pronunciation patterns and intonations and stores them in the database.

[0308] 3. Training of Deep Learning Model

[0309] The voice feature quantities are used to train the deep learning model. The server uses frameworks such as TensorFlow to learn specific ways of speaking and intonations. A large amount of voice data and the corresponding text data are required for this training. When the training is completed, the model is saved in a specific folder.

[0310] 4. Recognition of User's Emotion

[0311] When the user speaks to the smart speaker again with a question or conversation, the terminal records the voice again and transmits the voice data to the server. The server analyzes the received voice data and uses emotion recognition technology (e.g., IBM Watson (registered trademark) Tone Analyzer) to analyze the user's emotion. The analysis results are saved in JSON format.

[0312] 5. Analysis of the Intention of Questions and Conversations

[0313] The server converts the audio data back into text and performs morphological analysis. For example, MeCab is used for this purpose. Morphological analysis helps understand the user's intent and obtain the information necessary to generate an appropriate response.

[0314] 6. Response generation and adjustment

[0315] Using a pre-trained deep learning model, the server generates appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, considering the user's mood, it might generate additional follow-up responses such as, "You seem a little down, are you okay?" The generated text responses are then converted into speech data using the Google Text-to-Speech API.

[0316] 7. Lip sync and avatar display

[0317] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's facial expressions and mouth movements. The terminal receives this audio data and lip-sync data, playing the audio while simultaneously displaying the avatar's facial expressions and mouth movements on the screen.

[0318] 8. User experience

[0319] Users can hear voice responses from smart speakers and see the natural facial expressions and mouth movements of the avatar on the display. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[0320] Specific example

[0321] For example, if a user says "Good morning" to a smart speaker, the smart speaker records the voice and sends it to a server. The server analyzes this voice, extracts features, and stores them in a database. Then, it trains a deep learning model using TensorFlow. If the user asks another question, such as "How was your day?", the server uses emotion recognition technology to analyze the user's emotions. For example, if the user sounds sad, it might generate a response like, "I had a relaxing day. You seem a little down, are you okay?" and convert it into speech using the Google Text-to-Speech API. The device plays this speech, and the display shows the avatar's facial expressions and mouth movements.

[0322] Example of a prompt

[0323] Please generate example responses to the question, "How was your day?" Include additional, emotionally sensitive follow-up if the user seems sad.

[0324] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0325] Step 1: Collect and save audio samples

[0326] The user imitates a specific person's voice and says "Good morning" to the smart speaker. The smart speaker, as the device, records this voice using its microphone. The recorded voice data is temporarily stored in internal memory and then sent to a server via Wi-Fi.

[0327] Input: User's voice "Good morning"

[0328] Output: Recorded audio data (digital format)

[0329] Step 2: Analysis of the voice moon

[0330] The server stores the received audio data and uses speech recognition technology to convert the audio into text data. Using the Google Speech-to-Text API, the audio data is converted into the text "Good morning." Based on this text data, speech features such as pronunciation patterns and intonation are extracted and stored in a new database.

[0331] Input: Recorded audio data

[0332] Output: Text data and speech features

[0333] Step 3: Training the deep learning model

[0334] The server trains a deep learning model using speech features. Libraries such as TensorFlow are used to learn specific speaking styles and intonations. In this training example, a large amount of speech data and corresponding text data are used to enable the model to mimic specific speaking styles and intonations. The trained model is saved to a specific folder.

[0335] Input: Speech features

[0336] Output: Trained deep learning model

[0337] Step 4: User emotion recognition

[0338] The user speaks to the smart speaker again, for example, asking, "How was your day?" The device records this voice and sends it to the server. The server receives the voice data and analyzes the user's emotions using emotion recognition technology (e.g., IBM Watson Tone Analyzer). The analysis results are stored in JSON format.

[0339] Input: User's voice "How was your day?"

[0340] Output: User sentiment analysis results

[0341] Step 5: Analyzing the intent behind questions and conversations

[0342] The server then converts the received audio data back into text data and performs morphological analysis. This is done using a morphological analyzer such as MeCab. By converting the audio into text data such as "How was your day?" and performing morphological analysis, the server understands what the user is asking.

[0343] Input: Audio data

[0344] Output: Analyzed text data and user intent

[0345] Step 6: Generating and adjusting the response

[0346] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, based on the sentiment recognition results, it might generate a follow-up response such as, "You seem a little down, are you okay?" This generated response is then converted to speech using the Google Text-to-Speech API.

[0347] Input: Analyzed text data and user sentiment analysis results

[0348] Output: Adjusted voice response

[0349] Step 7: Lip sync and avatar display

[0350] The server generates lip-sync data that synchronizes with the generated audio data. This lip-sync data naturally expresses the avatar's mouth movements and facial expressions. The terminal receives the audio data and lip-sync data, plays the audio, and simultaneously displays the avatar on the screen. The avatar's mouth moves in accordance with the lip-sync data, displaying appropriate facial expressions.

[0351] Input: Adjusted voice response

[0352] Output: Avatar facial expressions corresponding to audio and lip-sync data

[0353] Step 8: User Experience

[0354] Users can hear voice responses played from the smart speaker. They can also feel like they are actually having a conversation by observing the natural mouth movements and facial expressions of the avatar displayed on the screen. This system is designed to allow users to enjoy a natural conversational experience with a specific person.

[0355] Input: User's question

[0356] Output: Natural conversation experience and avatar display

[0357] (Application Example 2)

[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0359] A system that can reproduce the voice of a specific person in a remote location and adjust its response by recognizing the user's emotions can provide a natural, emotionally resonant conversational experience in the home and in daily life. However, while such a system could also be applied to work environments such as factories to provide emotional support to employees and improve work efficiency, current systems do not adequately address this need. There is a need for a new system that allows factory workers to relax and boost morale through direct interaction.

[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice samples of a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, and means for recognizing the user's emotions and adjusting the response based on the results. This makes it possible to provide a natural conversational experience that is empathetic to emotions, even in work environments such as factories, enabling mental support for employees and improvement of work efficiency.

[0361] A "voice sample" refers to audio data recorded for use in other systems or algorithms, specifically recording the voice of a particular person.

[0362] "Speech features" refer to characteristic data representations extracted from speech data, such as pronunciation patterns and intonation.

[0363] A "deep learning model" refers to a type of machine learning model that is constructed using multi-layer neural networks to learn complex patterns from data.

[0364] "User voice input" refers to the act of a system user entering information by voice, or the input data itself.

[0365] "Response" refers to the reply or content of the response that the system generates in response to voice input from the user.

[0366] An "avatar" refers to a virtual person or character displayed on a computer screen or by a robot, whose mouth movements and facial expressions are synchronized with the voice.

[0367] "Emotion recognition" refers to a technology or process that analyzes audio data and other information to identify the emotional state of a speaker.

[0368] A "server" refers to a computer system that stores and processes data on a network and communicates with client devices.

[0369] "Means of collection" refers to hardware and software for recording or collecting audio samples via a network.

[0370] "Means of analysis" refers to algorithms and software used to analyze collected data and extract and process speech features.

[0371] "Training methods" refer to the process and techniques of optimizing deep learning models using speech features to improve their skills so that they can accurately generate responses.

[0372] "Means of output" refers to speakers or speech synthesis technology that allow the user to hear the generated response as audio.

[0373] "Means of display" refers to monitors and display technologies used to show users the mouth movements and facial expressions of avatars.

[0374] "Means of adjustment" refers to algorithms and technologies that adjust responses generated based on the results of emotion recognition to create appropriate responses that match the user's emotions.

[0375] This invention provides a system for use in work environments such as factories that reproduces the voice of a specific person, recognizes the emotions of employees, and generates appropriate responses. This can provide emotional support to employees and improve work efficiency.

[0376] The server stores collected audio samples and converts them into text data using speech recognition technology. It also extracts speech features such as pronunciation patterns and intonation, stores them in a database, and uses them to train a deep learning model. The server trains the deep learning model to create a model that learns specific speaking styles and intonations. Then, it uses the trained model to generate appropriate responses to user questions and outputs them as audio. For audio output, text-to-speech (TTS) technology is used to convert the text responses into audio data.

[0377] Furthermore, the server recognizes the user's emotions and adjusts the response generated based on that. For example, if the user appears sad, it generates a response that takes those emotions into consideration. In this way, a natural conversational experience is achieved.

[0378] The terminal (such as a robot or monitor) uses audio and lip-sync data received from the server to play the audio and display the avatar's facial expressions and mouth movements. This allows the user to hear the audio response and see the avatar's natural mouth movements and expressions.

[0379] Specifically, the `speech_recognition` library is used for speech recognition, and the `transformers` library is used for emotion recognition. The `pyttsx3` library is used to convert text to speech.

[0380] The following are specific examples of prompt statements.

[0381] "Please generate an appropriate response based on my emotions. Analyze the following text to recognize my emotions and then provide a friendly voice response accordingly."

[0382] In this way, the present invention provides a natural conversational experience that is attentive to the emotions of employees, even in work environments such as factories, thereby achieving emotional support and improved work efficiency.

[0383] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0384] Step 1:

[0385] The user provides a voice sample. The user speaks to the robot and provides a voice sample of a specific person. For example, the user might say "Good morning." This voice sample is recorded by the device and sent to the server.

[0386] Step 2:

[0387] The server analyzes the audio sample. The server stores the received audio data and converts it into text data using speech recognition technology (e.g., the speech_recognition library). Based on the converted text, it extracts speech features such as pronunciation patterns and intonation. These features are stored in a database.

[0388] Step 3:

[0389] The server trains a deep learning model. The server uses accumulated speech features to train a deep learning model to learn specific speech patterns and intonations. During this process, the transformers library is used to optimize the model using speech features as input. The trained model is then used to generate future user responses.

[0390] Step 4:

[0391] The system recognizes the user's emotions. When the user speaks to the robot again, the terminal records the voice again and sends it to the server. The server analyzes the received voice data and uses emotion recognition technology (e.g., an emotion recognition model from the transformers library) to analyze the user's emotions. The emotion label is output from the voice data as input.

[0392] Step 5:

[0393] The system analyzes the intent behind questions and conversations. The server converts the audio data into text (for example, using the speech_recognition library) and performs morphological analysis to understand the user's intent. This analysis outputs the intent in text format, and a response is determined based on it.

[0394] Step 6:

[0395] The server generates a response and outputs it as speech. A trained deep learning model is used to generate an appropriate response to the user's question. For example, a response such as "I had a relaxing day today" might be generated. The generated text response is converted into speech data using speech synthesis technology (e.g., the pyttsx3 library) and output.

[0396] Step 7:

[0397] The device performs lip-syncing and displays the avatar. The server generates lip-sync data corresponding to the audio data (for example, using an animation library) to create data that naturally expresses the avatar's mouth movements and facial expressions. The device uses this data to play the audio while displaying the avatar's mouth movements and facial expressions.

[0398] Step 8:

[0399] The user experience is enhanced. Users can listen to responses generated through the robot and see the avatar's natural mouth movements and facial expressions. This allows users to enjoy a natural conversational experience with family members or deceased loved ones who are far away, and receive responses that are more emotionally resonant.

[0400] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0401] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0402] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0403] [Second Embodiment]

[0404] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0405] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0406] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0407] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0408] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0410] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0411] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0412] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0413] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0414] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0415] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0416] The present invention mainly involves the following processing steps.

[0417] 1. Collection of audio samples:

[0418] Voice samples are collected when users speak to a smart speaker.

[0419] The device (smart speaker) records the user's voice and sends that data to the server.

[0420] For example, when a user says "Good morning," the voice is recorded by the smart speaker, and the voice data is sent to the server.

[0421] 2. Storage and analysis of audio data:

[0422] The server temporarily stores the received audio data.

[0423] The saved audio data is converted into text using speech recognition technology. Additionally, audio features (pronunciation patterns, intonation, etc.) are extracted.

[0424] The extracted speech features are stored in a database and used later for deep learning.

[0425] 3. Training deep learning models:

[0426] The server uses the accumulated audio features to train a deep learning model.

[0427] During the training process, specific speaking styles and intonations are learned, and models are built for generating natural dialogue.

[0428] For example, phrases that users repeatedly say, such as "Hello" or "How are you?", can be used as training data.

[0429] 4. Receiving and analyzing questions and conversations:

[0430] When a user speaks a question or engages in a conversation with a smart speaker, the device records it and sends it to a server.

[0431] The server analyzes the received audio, converts it to text, and performs morphological analysis to understand the user's intent.

[0432] For example, if a user asks "What did you do today?", the content of that question is converted into text and analyzed.

[0433] 5. Response generation and speech synthesis:

[0434] A trained deep learning model is used to generate appropriate responses to user questions.

[0435] The generated response is converted into speech data using text-to-speech (TTS) technology.

[0436] For example, a response like "I had a relaxing day today" might be generated.

[0437] 6. Lip sync and avatar display:

[0438] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0439] The terminal receives voice data and lip-sync data, and uses them to actually respond to the user.

[0440] The user can hear the response and simultaneously see the avatar's natural mouth movements.

[0441] For example, an avatar saying "I had a relaxing day today" will be displayed with accurate mouth movements.

[0442] Thus, the present invention provides a system that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away, by combining a technology that reproduces the voice characteristics of a specific person with a lip-sync technology that synchronizes with the voice. It also includes specific means to make the user feel psychologically secure and familiar with the system.

[0443] The following describes the processing flow.

[0444] Step 1:

[0445] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[0446] Step 2:

[0447] The device (smart speaker) records the user's voice and saves it to a temporary buffer. After recording is complete, it sends the voice data to the server.

[0448] Step 3:

[0449] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[0450] Step 4:

[0451] The server converts audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features, including pronunciation patterns and intonation.

[0452] Step 5:

[0453] The server stores the extracted audio features in a database. Simultaneously, it uses this data to train a deep learning model. A multi-layered neural network is used for training.

[0454] Step 6:

[0455] The user initiates a new conversation or asks a question to the smart speaker. For example, they might ask, "How was your day?"

[0456] Step 7:

[0457] The device (smart speaker) re-records the user's voice and sends it to the server.

[0458] Step 8:

[0459] The server receives the audio data and converts it into text data. It then performs morphological analysis on the text to analyze the user's intent.

[0460] Step 9:

[0461] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a response like, "I had a relaxing day."

[0462] Step 10:

[0463] The server converts the generated text response into speech data using text-to-speech (TTS) technology. This enables natural-sounding speech.

[0464] Step 11:

[0465] The server generates lip-sync data corresponding to the audio data. This allows the audio and the avatar's mouth movements to be synchronized.

[0466] Step 12:

[0467] The server sends audio data and lip-sync data to the terminal.

[0468] Step 13:

[0469] The device (smart speaker) receives audio data and lip-sync data from the server. Based on the received data, it plays the audio and simultaneously displays the mouth movements of the avatar.

[0470] Step 14:

[0471] Users listen to responses generated through smart speakers and observe the natural mouth movements of their avatars. This allows them to enjoy a natural conversational experience with loved ones.

[0472] In this way, a system is created that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away.

[0473] (Example 1)

[0474] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0475] Conventional speech recognition and dialogue generation systems struggle to reproduce the natural speaking style and intonation of a specific person located remotely. As a result, they are insufficient as systems that provide users with a sense of psychological security and familiarity. Furthermore, these systems often lack visual feedback for the generated speech, leading to mismatches between speech and mouth movements. This, in turn, negatively impacts the user experience.

[0476] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0477] In this invention, the server includes means for acquiring voice data, means for storing the acquired voice data, means for converting the stored voice data into text using speech recognition technology, means for extracting and accumulating voice features, means for training a deep learning model using the accumulated voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice using speech synthesis technology, and means for displaying the mouth movements of an avatar in synchronization with the output voice. This makes it possible to generate a response in which the voice and mouth movements are naturally synchronized, providing the user with a sense of psychological security and familiarity.

[0478] "Means of acquisition" refers to devices or mechanisms for collecting the user's voice.

[0479] "Means of storage" refers to a storage device or service for temporarily storing the collected audio data.

[0480] "Speech recognition technology" refers to algorithms and software used to convert speech data into text data.

[0481] "Speech features" are identifiable characteristics such as pronunciation patterns and intonation that are extracted from speech data.

[0482] "Means of storage" refer to databases or storage devices for saving and managing extracted speech features.

[0483] A "deep learning model" is a neural network model that is trained using speech features.

[0484] "Means for generating responses in response to voice input" refers to a mechanism that uses a trained deep learning model to create appropriate responses to user voice input.

[0485] "Speech synthesis technology" refers to algorithms and software used to convert generated text data into speech data.

[0486] "Means for displaying the mouth movements of an avatar" refers to a display device or software that reproduces the mouth movements of an avatar in synchronization with the outputted audio.

[0487] To implement this invention, three elements are necessary: ​​a user, a terminal (smart speaker), and a server. The user speaks to the terminal, the terminal sends voice data to the server, and the server analyzes and processes the voice data to generate an appropriate response.

[0488] When a user speaks into a smart speaker, the device uses its microphone to capture the user's voice. Specifically, the user might say things like "Good morning" or "What did you do today?". This voice data is captured by the device and sent to a server via the internet.

[0489] The transmitted audio data is first temporarily stored on the server. The server converts the stored audio data into text using speech recognition technology. Services such as the Google Speech-to-Text API are used for this process. The text data obtained through speech recognition technology is stored in a database on the server.

[0490] Next, the server extracts speech features (e.g., pronunciation patterns, intonation) from the audio data. These speech features are used as data to train a deep learning model (e.g., TensorFlow). During the training process, the server learns the user's unique speaking style and intonation, and builds a model for generating natural-sounding dialogue.

[0491] When the user speaks a question or has a conversation into the device again, the device records the audio and sends it to the server. The server converts the audio back into text and performs morphological analysis (e.g., using MeCab). Based on the analyzed text data, the server uses a trained deep learning model to generate an appropriate response. For example, it might generate a response such as, "I had a relaxing day."

[0492] The generated text responses are converted into audio data using speech synthesis technology. Speech synthesis services such as the Google Text-to-Speech API are used for this process. The server then generates lip-sync data synchronized with the audio data to recreate the avatar's mouth movements.

[0493] The device uses the received audio and lip-sync data to respond to the user. The user can listen to the avatar's response, such as "I had a relaxing day," while observing the avatar's natural mouth movements.

[0494] Specific example

[0495] One day, a user asks a smart speaker, "What's the weather like tomorrow?" The device records the voice and sends it to a server. The server converts the voice to text, performs further analysis, and generates an appropriate response. This response is "It will be sunny tomorrow," and is converted into voice data via speech synthesis technology. The generated voice data and lip-sync data are sent to the device, and an avatar speaks "It will be sunny tomorrow," with its mouth movements naturally reproduced.

[0496] Example of a prompt

[0497] 1. "What's the weather like tomorrow?"

[0498] 2. "Please tell me about the latest news."

[0499] 3. "What fun things happened to you today?"

[0500] This system generates responses where voice and mouth movements are naturally synchronized, providing users with a sense of psychological reassurance and familiarity.

[0501] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0502] Step 1:

[0503] The user speaks to the smart speaker.

[0504] Input: Voice (Example: "Good morning")

[0505] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0506] Output: Recorded audio data

[0507] Step 2:

[0508] The device (smart speaker) sends the recorded audio data to the server.

[0509] Input: Recorded audio data

[0510] Specific operation: Captured audio data is sent to the server via the internet.

[0511] Output: Audio data sent to the server

[0512] Step 3:

[0513] The server temporarily stores the received audio data.

[0514] Input: Audio data sent to the server

[0515] Specific operation: The audio data is saved to storage (e.g., Amazon S3).

[0516] Output: Saved audio data

[0517] Step 4:

[0518] The server uses speech recognition technology to convert the audio data into text.

[0519] Input: Saved audio data

[0520] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and text data is returned.

[0521] Output: Text data

[0522] Step 5:

[0523] The server extracts speech features from the audio data and stores them in a database.

[0524] Input: Saved audio data

[0525] Specific operation: Audio data is analyzed, features such as pronunciation patterns and intonation are extracted, and these are stored in a database.

[0526] Output: Speech feature data

[0527] Step 6:

[0528] The server uses speech features to train a deep learning model.

[0529] Input: Speech feature data

[0530] Specific operation: A neural network is trained using a framework such as TensorFlow with audio feature data.

[0531] Output: Trained deep learning model

[0532] Step 7:

[0533] The user speaks to the smart speaker, asking questions or engaging in conversation.

[0534] Input: Voice (Example: "What did you do today?")

[0535] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0536] Output: Recorded audio data

[0537] Step 8:

[0538] The device sends the recorded audio data to the server.

[0539] Input: Recorded audio data

[0540] Specific operation: Captured audio data is sent to the server via the internet.

[0541] Output: Audio data sent to the server

[0542] Step 9:

[0543] The server converts the audio data into text and performs morphological analysis.

[0544] Input: Audio data sent to the server

[0545] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and after obtaining text data, morphological analysis is performed using MeCab, etc.

[0546] Output: Morphologically analyzed text data

[0547] Step 10:

[0548] The server uses a trained deep learning model to generate appropriate responses to questions and conversations.

[0549] Input: Morphologically analyzed text data

[0550] Specific operation: Input text data into a trained deep learning model and generate an appropriate response.

[0551] Output: Generated text response

[0552] Step 11:

[0553] The server converts the generated text response into speech data using speech synthesis technology.

[0554] Input: Generated text response

[0555] Specific operation: Convert text responses into speech data using the Google Text-to-Speech API, etc.

[0556] Output: Generated audio data

[0557] Step 12:

[0558] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0559] Input: Generated audio data

[0560] Specific operation: Analyzes audio data and creates lip-sync data to match the audio.

[0561] Output: Lip-sync data

[0562] Step 13:

[0563] The device receives voice data and lip-sync data and displays a response to the user.

[0564] Input: Generated audio data and lip-sync data

[0565] Specific operation: Based on the received data, the avatar responds to the user with voice and simultaneously displays the movement of its mouth.

[0566] Output: Voice responses and mouth movements by the avatar

[0567] In this way, through each processing step, users can enjoy natural conversations with specific individuals in remote locations through the reproduction of natural dialogue and the accompanying mouth movements of their avatars.

[0568] (Application Example 1)

[0569] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0570] Conventional speech recognition and voice response systems have limitations in achieving natural dialogue with users, and systems combining speech synthesis and lip-syncing have not been widely implemented. Furthermore, in food delivery services, it has been difficult to provide personalized menu suggestions and answer questions through dialogue with users. This invention aims to solve these problems and enable users to use food delivery services more comfortably.

[0571] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0572] In this invention, the server includes means for collecting voice samples from a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, virtual assistant means for analyzing the user's questions, generating an appropriate response, and outputting that response using speech synthesis technology, and means for using a generative AI model that generates a response based on the analysis results of the user's voice input. This enables a more natural and intuitive conversation when a user uses a food delivery service.

[0573] A "voice sample" is data that is a recording of a person speaking in a remote location.

[0574] "Speech features" refer to characteristic information such as pronunciation patterns and intonation extracted from collected speech samples.

[0575] A "deep learning model" is a model that uses a multi-layered neural network to learn from large amounts of data and perform predictions and classifications on new data.

[0576] A "virtual assistant" is a software system that generates appropriate responses to user voice input and outputs those responses using speech synthesis technology.

[0577] A "generative AI model" is an artificial intelligence model trained to generate responses based on a user's voice input.

[0578] A "lip-syncing method" is a means of displaying the mouth movements of an avatar in sync with the generated audio.

[0579] A "food delivery service" is a service that delivers food to customers based on their orders.

[0580] This invention is a system that collects voice samples from people in remote locations and uses them to provide users with natural, real-time conversations. This system is particularly effective in food delivery services, where users can use devices such as smartphones or smart glasses to interact with a virtual assistant and obtain information about menus and recommended dishes.

[0581] First, the server collects voice samples and temporarily stores the data. These voice samples are obtained when the user speaks into their smartphone or smart glasses. For example, if the user asks, "What pasta dish do you recommend?", the voice is recorded by the device and sent to the server.

[0582] Next, the server analyzes the collected audio data and converts it into text using speech recognition technology. Simultaneously, it extracts audio features (such as pronunciation patterns and intonation) and stores them in a database. This allows a deep learning model to be trained using the audio features.

[0583] A trained deep learning model generates appropriate responses to voice input from the user. When the user provides voice input again, the server analyzes the speech to understand the intent and uses the generative AI model to generate an appropriate response. This response is then converted into audio data using speech synthesis technology and output to the user.

[0584] Furthermore, lip-sync data synchronized with this generated audio data is created, allowing the virtual assistant avatar to reproduce natural mouth movements. This allows users to visually confirm that the virtual assistant is speaking naturally.

[0585] As a concrete example, when a user asks, "What pasta dish do you recommend?", the system generates a response through the following process: It replies with voice, "My recommendation is carbonara," and an avatar makes the same response with natural mouth movements. This allows the user to experience an immersive conversation.

[0586] The hardware used will be a smartphone, smart glasses, and a server. The software used will be speech recognition technology (e.g., the speech_recognition library), speech synthesis technology (e.g., the gtts library), and a generative AI model (e.g., the GPT-2 model from the transformers library).

[0587] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0588] Step 1:

[0589] When a user speaks to their smartphone or smart glasses, their voice is recorded by the device. The recorded voice data is then sent directly to the server. For example, if a user asks, "What pasta dish do you recommend?", that voice will be collected. The input is the user's voice data, and the output is the voice data sent to the server.

[0590] Step 2:

[0591] The server temporarily stores the received audio data. Then, it converts the audio data into text data using speech recognition technology. Here, the `speech_recognition` library is used to convert speech to text. The input is audio data, and the output is text data.

[0592] Step 3:

[0593] The server further analyzes the text data and extracts speech features (pronunciation patterns, intonation, etc.). A speech feature extraction algorithm is used for the analysis, and the extracted features are stored in a database. The input is text data, and the output is speech feature data.

[0594] Step 4:

[0595] The server accumulates speech feature data and uses it to train a deep learning model. A large amount of speech features are used for training, and a trained generative AI model is constructed. The input is speech feature data, and the output is the trained generative AI model.

[0596] Step 5:

[0597] When the user speaks a question to the device again, the device records the audio and sends it to the server. The server analyzes the received audio, converts it into text data, and performs morphological analysis to understand the user's intent. The input is the newly collected audio data, and the output is the analyzed text data and the results of the intent analysis.

[0598] Step 6:

[0599] The server uses a trained generative AI model to generate appropriate responses to user voice input. The generated responses are converted into speech data using speech synthesis technology (e.g., the GTTS library). The input is the intent analysis result, and the output is the speech response data.

[0600] Step 7:

[0601] The server generates lip-sync data synchronized with the generated audio data, ensuring that the virtual assistant avatar reflects natural mouth movements. The lip-sync data is used to control the avatar's mouth movements in real time. The input is audio data, and the output is lip-sync data.

[0602] Step 8:

[0603] The terminal receives audio data and lip-sync data from the server, plays an audio response to the user, and displays the avatar's lip-sync. The input is audio data and lip-sync data, and the output is the audio response to the user and the avatar's natural mouth movements.

[0604] Through the steps outlined above, users can ask questions about the food delivery service in a natural, conversational format. This allows users to obtain detailed information through dialogue, significantly improving convenience.

[0605] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0606] This invention is a system that reproduces the voice characteristics of a specific person in a remote location, recognizes the user's emotions, and adjusts the response accordingly. The system consists of a smart speaker and a server, and each process is performed as follows.

[0607] 1. Collection and storage of audio samples:

[0608] The user speaks to the smart speaker, which provides a voice sample of a specific person. For example, the user might say "Good morning."

[0609] The device (smart speaker) records the user's voice and sends that voice data to the server.

[0610] 2. Analysis of audio data:

[0611] The server stores the received audio data and converts it into text data using speech recognition technology. Based on the converted text, it extracts audio features such as pronunciation patterns and intonation.

[0612] The extracted speech features are stored in a database and used to train a deep learning model.

[0613] 3. Training deep learning models:

[0614] The server uses the accumulated speech features to train a deep learning model to learn specific speaking styles and intonations.

[0615] The trained model will be used to generate future user responses.

[0616] 4. User emotion recognition:

[0617] When a user speaks a new question or engages in a conversation with the smart speaker, the device records the audio again and sends it to the server.

[0618] The server analyzes the received audio data and uses emotion recognition technology to analyze the user's emotions.

[0619] 5. Analyzing the intent behind questions and conversations:

[0620] The server converts the audio data into text, performs morphological analysis, and understands the user's intent.

[0621] 6. Response generation and adjustment:

[0622] Using a pre-trained deep learning model, it generates appropriate responses to user questions. For example, it can generate a response like, "I had a relaxing day today."

[0623] The response is adjusted based on the user's emotions as assessed by the emotion recognition engine. For example, if the user appears sad, it might generate a response such as, "I had a relaxing day today. You seem a little down, are you okay?"

[0624] The generated text responses are converted into speech data using text-to-speech (TTS) technology.

[0625] 7. Lip sync and avatar display:

[0626] The server generates lip-sync data corresponding to the audio data, naturally expressing the avatar's mouth movements and facial expressions.

[0627] The device receives audio data and lip-sync data, plays the audio, and displays the avatar's facial expressions and mouth movements.

[0628] 8. User experience:

[0629] Users can listen to the responses generated through the smart speaker and see the avatar's natural mouth movements and facial expressions.

[0630] This allows users to enjoy natural conversations with family members or deceased loved ones who live far away, and to receive responses that are sensitive to the user's emotions.

[0631] In this way, the present invention realizes a system that realistically reproduces the voice characteristics of a specific person, understands the user's emotions, and provides a more natural and approachable conversational experience.

[0632] The following describes the processing flow.

[0633] Step 1:

[0634] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[0635] Step 2:

[0636] The device (smart speaker) records the user's voice and saves it to a temporary buffer. It then sends the recorded voice data to the server.

[0637] Step 3:

[0638] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[0639] Step 4:

[0640] The server converts the audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features (such as pronunciation patterns and intonation).

[0641] Step 5:

[0642] The server stores the extracted speech features in a database. Simultaneously, this data is used to train a deep learning model.

[0643] Step 6:

[0644] The user initiates a new question or conversation with the smart speaker. For example, they might ask, "How was your day?"

[0645] Step 7:

[0646] The device (smart speaker) re-records the user's voice and sends it to the server.

[0647] Step 8:

[0648] The server analyzes the received audio data and converts it back into text data. This text is then subjected to morphological analysis to determine the user's intent.

[0649] Step 9:

[0650] The server uses an emotion recognition engine during the process of analyzing voice data to analyze the user's emotions (such as joy, anger, sadness, etc.).

[0651] Step 10:

[0652] The server evaluates the user's emotions and uses a trained deep learning model to generate an appropriate response. For example, it might generate a response like, "I had a relaxing day."

[0653] Step 11:

[0654] The server adjusts its response based on the emotion recognition results. For example, if the user seems sad, it might change the response to something like, "I had a relaxing day today. You seem a little down, are you okay?"

[0655] Step 12:

[0656] The generated text response is converted into speech data using text-to-speech (TTS) technology. This makes the response sound more natural.

[0657] Step 13:

[0658] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's mouth movements and facial expressions.

[0659] Step 14:

[0660] The server sends audio data and lip-sync data to the terminal.

[0661] Step 15:

[0662] The device (smart speaker) receives audio data and lip-sync data from the server. It plays the audio and displays the avatar's facial expressions and mouth movements.

[0663] Step 16:

[0664] Users listen to the generated responses and observe the avatar's natural mouth movements and facial expressions. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[0665] This specific processing flow allows the present invention to faithfully reproduce the voice characteristics and pronunciation patterns of a particular person and provide a friendly response that corresponds to the user's emotions.

[0666] (Example 2)

[0667] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0668] Providing a natural and engaging conversational experience with a specific person located remotely is challenging. Furthermore, recognizing the user's emotions and generating appropriate responses accordingly is also difficult. Therefore, the challenge lies in accurately reproducing a specific person's voice and speaking style while simultaneously providing responses that resonate with the user's emotions.

[0669] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0670] In this invention, the server includes means for collecting voice samples of a specific person located remotely, means for analyzing the collected voice samples to extract voice features such as pronunciation patterns and intonation, means for training a deep learning model using the voice features, means for recognizing the user's emotions and adjusting the response content, and means for converting voice input from the user into text and analyzing the intent. This makes it possible to realistically reproduce the voice of a specific person and provide a natural conversational experience that responds to the user's emotions.

[0671] A "voice sample" is a digital recording of a user's speech, which imitates the voice of a specific person.

[0672] "Pronunciation patterns" refer to characteristics of speech such as the flow of phonemes, rhythm, stress, and accentuation.

[0673] "Intonation" refers to the changes in pitch and intonation of a voice in spoken language.

[0674] "Speech features" are specific data points extracted from speech samples, and include elements such as pronunciation patterns, intonation, pitch, and length.

[0675] A "deep learning model" is a type of artificial intelligence trained using large amounts of data, and specifically refers to models that use neural networks.

[0676] "Emotion recognition" refers to technology that analyzes and identifies a user's emotional state from their voice or text.

[0677] An "avatar" refers to a representation of a person or character displayed in a virtual space for interaction with a user.

[0678] "Lip sync" refers to a technology that synchronizes the mouth movements of an avatar with the audio.

[0679] "Speech recognition technology" is a technology for converting speech into text, and includes the process of analyzing a speech sample and converting it into text data.

[0680] "Speech synthesis technology" refers to the technology that generates speech based on text data.

[0681] Modes for carrying out the invention

[0682] This invention relates to a system that reproduces the voice characteristics of a specific person located remotely, recognizes the user's emotions, and adjusts its response accordingly. The system consists of a smart speaker and a server and includes the following elements:

[0683] 1. Collection and storage of audio samples

[0684] Voice samples are collected when the user speaks to the smart speaker. For example, the user says "Good morning." The device (smart speaker) records the user's voice with its microphone and temporarily stores this voice data in its internal memory. Next, the device sends this voice data to the server via Wi-Fi.

[0685] 2. Analysis of audio data

[0686] The server stores the received audio data in a dedicated database. Then, it uses speech recognition technology to convert the audio data into text data. Specifically, it uses the Google Speech-to-Text API to convert the audio to text. Based on this text data, it extracts speech features such as pronunciation patterns and intonation, and stores them in the database.

[0687] 3. Training of deep learning models

[0688] A deep learning model is trained using speech features. The server uses a framework such as TensorFlow to learn specific speech patterns and intonations. This training requires a large amount of speech data and corresponding text data. Once training is complete, the model is saved to a specific folder.

[0689] 4. User emotion recognition

[0690] When the user speaks a question or engages in conversation with the smart speaker again, the device records the voice again and sends the audio data to the server. The server analyzes the received audio data and uses emotion recognition technology (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. This analysis result is stored in JSON format.

[0691] 5. Analyzing the intent behind questions and conversations

[0692] The server converts the audio data back into text and performs morphological analysis. For example, MeCab is used for this purpose. Morphological analysis helps understand the user's intent and obtain the information necessary to generate an appropriate response.

[0693] 6. Response generation and adjustment

[0694] Using a pre-trained deep learning model, the server generates appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, considering the user's mood, it might generate additional follow-up responses such as, "You seem a little down, are you okay?" The generated text responses are then converted into speech data using the Google Text-to-Speech API.

[0695] 7. Lip sync and avatar display

[0696] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's facial expressions and mouth movements. The terminal receives this audio data and lip-sync data, playing the audio while simultaneously displaying the avatar's facial expressions and mouth movements on the screen.

[0697] 8. User experience

[0698] Users can hear voice responses from smart speakers and see the natural facial expressions and mouth movements of the avatar on the display. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[0699] Specific example

[0700] For example, if a user says "Good morning" to a smart speaker, the smart speaker records the voice and sends it to a server. The server analyzes this voice, extracts features, and stores them in a database. Then, it trains a deep learning model using TensorFlow. If the user asks another question, such as "How was your day?", the server uses emotion recognition technology to analyze the user's emotions. For example, if the user sounds sad, it might generate a response like, "I had a relaxing day. You seem a little down, are you okay?" and convert it into speech using the Google Text-to-Speech API. The device plays this speech, and the display shows the avatar's facial expressions and mouth movements.

[0701] Example of a prompt

[0702] Please generate example responses to the question, "How was your day?" Include additional, emotionally sensitive follow-up if the user seems sad.

[0703] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0704] Step 1: Collect and save audio samples

[0705] The user imitates a specific person's voice and says "Good morning" to the smart speaker. The smart speaker, as the device, records this voice using its microphone. The recorded voice data is temporarily stored in internal memory and then sent to a server via Wi-Fi.

[0706] Input: User's voice "Good morning"

[0707] Output: Recorded audio data (digital format)

[0708] Step 2: Analysis of the voice moon

[0709] The server stores the received audio data and uses speech recognition technology to convert the audio into text data. Using the Google Speech-to-Text API, the audio data is converted into the text "Good morning." Based on this text data, speech features such as pronunciation patterns and intonation are extracted and stored in a new database.

[0710] Input: Recorded audio data

[0711] Output: Text data and speech features

[0712] Step 3: Training the deep learning model

[0713] The server trains a deep learning model using speech features. Libraries such as TensorFlow are used to learn specific speaking styles and intonations. In this training example, a large amount of speech data and corresponding text data are used to enable the model to mimic specific speaking styles and intonations. The trained model is saved to a specific folder.

[0714] Input: Speech features

[0715] Output: Trained deep learning model

[0716] Step 4: User emotion recognition

[0717] The user speaks to the smart speaker again, for example, asking, "How was your day?" The device records this voice and sends it to the server. The server receives the voice data and analyzes the user's emotions using emotion recognition technology (e.g., IBM Watson Tone Analyzer). The analysis results are stored in JSON format.

[0718] Input: User's voice "How was your day?"

[0719] Output: User sentiment analysis results

[0720] Step 5: Analyzing the intent behind questions and conversations

[0721] The server then converts the received audio data back into text data and performs morphological analysis. This is done using a morphological analyzer such as MeCab. By converting the audio into text data such as "How was your day?" and performing morphological analysis, the server understands what the user is asking.

[0722] Input: Audio data

[0723] Output: Analyzed text data and user intent

[0724] Step 6: Generating and adjusting the response

[0725] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, based on the sentiment recognition results, it might generate a follow-up response such as, "You seem a little down, are you okay?" This generated response is then converted to speech using the Google Text-to-Speech API.

[0726] Input: Analyzed text data and user sentiment analysis results

[0727] Output: Adjusted voice response

[0728] Step 7: Lip sync and avatar display

[0729] The server generates lip-sync data that synchronizes with the generated audio data. This lip-sync data naturally expresses the avatar's mouth movements and facial expressions. The terminal receives the audio data and lip-sync data, plays the audio, and simultaneously displays the avatar on the screen. The avatar's mouth moves in accordance with the lip-sync data, displaying appropriate facial expressions.

[0730] Input: Adjusted voice response

[0731] Output: Avatar facial expressions corresponding to audio and lip-sync data

[0732] Step 8: User Experience

[0733] Users can hear voice responses played from the smart speaker. They can also feel like they are actually having a conversation by observing the natural mouth movements and facial expressions of the avatar displayed on the screen. This system is designed to allow users to enjoy a natural conversational experience with a specific person.

[0734] Input: User's question

[0735] Output: Natural conversation experience and avatar display

[0736] (Application Example 2)

[0737] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0738] A system that can reproduce the voice of a specific person in a remote location and adjust its response by recognizing the user's emotions can provide a natural, emotionally resonant conversational experience in the home and in daily life. However, while such a system could also be applied to work environments such as factories to provide emotional support to employees and improve work efficiency, current systems do not adequately address this need. There is a need for a new system that allows factory workers to relax and boost morale through direct interaction.

[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice samples of a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, and means for recognizing the user's emotions and adjusting the response based on the results. This makes it possible to provide a natural conversational experience that is empathetic to emotions, even in work environments such as factories, enabling mental support for employees and improvement of work efficiency.

[0740] A "voice sample" refers to audio data recorded for use in other systems or algorithms, specifically recording the voice of a particular person.

[0741] "Speech features" refer to characteristic data representations extracted from speech data, such as pronunciation patterns and intonation.

[0742] A "deep learning model" refers to a type of machine learning model that is constructed using multi-layer neural networks to learn complex patterns from data.

[0743] "User voice input" refers to the act of a system user entering information by voice, or the input data itself.

[0744] "Response" refers to the reply or content of the response that the system generates in response to voice input from the user.

[0745] An "avatar" refers to a virtual person or character displayed on a computer screen or by a robot, whose mouth movements and facial expressions are synchronized with the voice.

[0746] "Emotion recognition" refers to a technology or process that analyzes audio data and other information to identify the emotional state of a speaker.

[0747] A "server" refers to a computer system that stores and processes data on a network and communicates with client devices.

[0748] "Means of collection" refers to hardware and software for recording or collecting audio samples via a network.

[0749] "Means of analysis" refers to algorithms and software used to analyze collected data and extract and process speech features.

[0750] "Training methods" refer to the process and techniques of optimizing deep learning models using speech features to improve their skills so that they can accurately generate responses.

[0751] "Means of output" refers to speakers or speech synthesis technology that allow the user to hear the generated response as audio.

[0752] "Means of display" refers to monitors and display technologies used to show users the mouth movements and facial expressions of avatars.

[0753] "Means of adjustment" refers to algorithms and technologies that adjust responses generated based on the results of emotion recognition to create appropriate responses that match the user's emotions.

[0754] This invention provides a system for use in work environments such as factories that reproduces the voice of a specific person, recognizes the emotions of employees, and generates appropriate responses. This can provide emotional support to employees and improve work efficiency.

[0755] The server stores collected audio samples and converts them into text data using speech recognition technology. It also extracts speech features such as pronunciation patterns and intonation, stores them in a database, and uses them to train a deep learning model. The server trains the deep learning model to create a model that learns specific speaking styles and intonations. Then, it uses the trained model to generate appropriate responses to user questions and outputs them as audio. For audio output, text-to-speech (TTS) technology is used to convert the text responses into audio data.

[0756] Furthermore, the server recognizes the user's emotions and adjusts the response generated based on that. For example, if the user appears sad, it generates a response that takes those emotions into consideration. In this way, a natural conversational experience is achieved.

[0757] The terminal (such as a robot or monitor) uses audio and lip-sync data received from the server to play the audio and display the avatar's facial expressions and mouth movements. This allows the user to hear the audio response and see the avatar's natural mouth movements and expressions.

[0758] Specifically, the `speech_recognition` library is used for speech recognition, and the `transformers` library is used for emotion recognition. The `pyttsx3` library is used to convert text to speech.

[0759] The following are specific examples of prompt statements.

[0760] "Please generate an appropriate response based on my emotions. Analyze the following text to recognize my emotions and then provide a friendly voice response accordingly."

[0761] In this way, the present invention provides a natural conversational experience that is attentive to the emotions of employees, even in work environments such as factories, thereby achieving emotional support and improved work efficiency.

[0762] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0763] Step 1:

[0764] The user provides a voice sample. The user speaks to the robot and provides a voice sample of a specific person. For example, the user might say "Good morning." This voice sample is recorded by the device and sent to the server.

[0765] Step 2:

[0766] The server analyzes the audio sample. The server stores the received audio data and converts it into text data using speech recognition technology (e.g., the speech_recognition library). Based on the converted text, it extracts speech features such as pronunciation patterns and intonation. These features are stored in a database.

[0767] Step 3:

[0768] The server trains a deep learning model. The server uses accumulated speech features to train a deep learning model to learn specific speech patterns and intonations. During this process, the transformers library is used to optimize the model using speech features as input. The trained model is then used to generate future user responses.

[0769] Step 4:

[0770] The system recognizes the user's emotions. When the user speaks to the robot again, the terminal records the voice again and sends it to the server. The server analyzes the received voice data and uses emotion recognition technology (e.g., an emotion recognition model from the transformers library) to analyze the user's emotions. The emotion label is output from the voice data as input.

[0771] Step 5:

[0772] The system analyzes the intent behind questions and conversations. The server converts the audio data into text (for example, using the speech_recognition library) and performs morphological analysis to understand the user's intent. This analysis outputs the intent in text format, and a response is determined based on it.

[0773] Step 6:

[0774] The server generates a response and outputs it as speech. A trained deep learning model is used to generate an appropriate response to the user's question. For example, a response such as "I had a relaxing day today" might be generated. The generated text response is converted into speech data using speech synthesis technology (e.g., the pyttsx3 library) and output.

[0775] Step 7:

[0776] The device performs lip-syncing and displays the avatar. The server generates lip-sync data corresponding to the audio data (for example, using an animation library) to create data that naturally expresses the avatar's mouth movements and facial expressions. The device uses this data to play the audio while displaying the avatar's mouth movements and facial expressions.

[0777] Step 8:

[0778] The user experience is enhanced. Users can listen to responses generated through the robot and see the avatar's natural mouth movements and facial expressions. This allows users to enjoy a natural conversational experience with family members or deceased loved ones who are far away, and receive responses that are more emotionally resonant.

[0779] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0780] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0781] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0782] [Third Embodiment]

[0783] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0784] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0785] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0786] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0787] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0788] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0789] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0790] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0791] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0792] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0793] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0794] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0795] The present invention mainly involves the following processing steps.

[0796] 1. Collection of audio samples:

[0797] Voice samples are collected when users speak to a smart speaker.

[0798] The device (smart speaker) records the user's voice and sends that data to the server.

[0799] For example, when a user says "Good morning," the voice is recorded by the smart speaker, and the voice data is sent to the server.

[0800] 2. Storage and analysis of audio data:

[0801] The server temporarily stores the received audio data.

[0802] The saved audio data is converted into text using speech recognition technology. Additionally, audio features (pronunciation patterns, intonation, etc.) are extracted.

[0803] The extracted speech features are stored in a database and used later for deep learning.

[0804] 3. Training deep learning models:

[0805] The server uses the accumulated audio features to train a deep learning model.

[0806] During the training process, specific speaking styles and intonations are learned, and models are built for generating natural dialogue.

[0807] For example, phrases that users repeatedly say, such as "Hello" or "How are you?", can be used as training data.

[0808] 4. Receiving and analyzing questions and conversations:

[0809] When a user speaks a question or engages in a conversation with a smart speaker, the device records it and sends it to a server.

[0810] The server analyzes the received audio, converts it to text, and performs morphological analysis to understand the user's intent.

[0811] For example, if a user asks "What did you do today?", the content of that question is converted into text and analyzed.

[0812] 5. Response generation and speech synthesis:

[0813] A trained deep learning model is used to generate appropriate responses to user questions.

[0814] The generated response is converted into speech data using text-to-speech (TTS) technology.

[0815] For example, a response like "I had a relaxing day today" might be generated.

[0816] 6. Lip sync and avatar display:

[0817] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0818] The terminal receives voice data and lip-sync data, and uses them to actually respond to the user.

[0819] The user can hear the response and simultaneously see the avatar's natural mouth movements.

[0820] For example, an avatar saying "I had a relaxing day today" will be displayed with accurate mouth movements.

[0821] Thus, the present invention provides a system that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away, by combining a technology that reproduces the voice characteristics of a specific person with a lip-sync technology that synchronizes with the voice. It also includes specific means to make the user feel psychologically secure and familiar with the system.

[0822] The following describes the processing flow.

[0823] Step 1:

[0824] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[0825] Step 2:

[0826] The device (smart speaker) records the user's voice and saves it to a temporary buffer. After recording is complete, it sends the voice data to the server.

[0827] Step 3:

[0828] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[0829] Step 4:

[0830] The server converts audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features, including pronunciation patterns and intonation.

[0831] Step 5:

[0832] The server stores the extracted audio features in a database. Simultaneously, it uses this data to train a deep learning model. A multi-layered neural network is used for training.

[0833] Step 6:

[0834] The user initiates a new conversation or asks a question to the smart speaker. For example, they might ask, "How was your day?"

[0835] Step 7:

[0836] The device (smart speaker) re-records the user's voice and sends it to the server.

[0837] Step 8:

[0838] The server receives the audio data and converts it into text data. It then performs morphological analysis on the text to analyze the user's intent.

[0839] Step 9:

[0840] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a response like, "I had a relaxing day."

[0841] Step 10:

[0842] The server converts the generated text response into speech data using text-to-speech (TTS) technology. This enables natural-sounding speech.

[0843] Step 11:

[0844] The server generates lip-sync data corresponding to the audio data. This allows the audio and the avatar's mouth movements to be synchronized.

[0845] Step 12:

[0846] The server sends audio data and lip-sync data to the terminal.

[0847] Step 13:

[0848] The device (smart speaker) receives audio data and lip-sync data from the server. Based on the received data, it plays the audio and simultaneously displays the mouth movements of the avatar.

[0849] Step 14:

[0850] Users listen to responses generated through smart speakers and observe the natural mouth movements of their avatars. This allows them to enjoy a natural conversational experience with loved ones.

[0851] In this way, a system is created that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away.

[0852] (Example 1)

[0853] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0854] Conventional speech recognition and dialogue generation systems struggle to reproduce the natural speaking style and intonation of a specific person located remotely. As a result, they are insufficient as systems that provide users with a sense of psychological security and familiarity. Furthermore, these systems often lack visual feedback for the generated speech, leading to mismatches between speech and mouth movements. This, in turn, negatively impacts the user experience.

[0855] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0856] In this invention, the server includes means for acquiring voice data, means for storing the acquired voice data, means for converting the stored voice data into text using speech recognition technology, means for extracting and accumulating voice features, means for training a deep learning model using the accumulated voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice using speech synthesis technology, and means for displaying the mouth movements of an avatar in synchronization with the output voice. This makes it possible to generate a response in which the voice and mouth movements are naturally synchronized, providing the user with a sense of psychological security and familiarity.

[0857] "Means of acquisition" refers to devices or mechanisms for collecting the user's voice.

[0858] "Means of storage" refers to a storage device or service for temporarily storing the collected audio data.

[0859] "Speech recognition technology" refers to algorithms and software used to convert speech data into text data.

[0860] "Speech features" are identifiable characteristics such as pronunciation patterns and intonation that are extracted from speech data.

[0861] "Means of storage" refer to databases or storage devices for saving and managing extracted speech features.

[0862] A "deep learning model" is a neural network model that is trained using speech features.

[0863] "Means for generating responses in response to voice input" refers to a mechanism that uses a trained deep learning model to create appropriate responses to user voice input.

[0864] "Speech synthesis technology" refers to algorithms and software used to convert generated text data into speech data.

[0865] "Means for displaying the mouth movements of an avatar" refers to a display device or software that reproduces the mouth movements of an avatar in synchronization with the outputted audio.

[0866] To implement this invention, three elements are necessary: ​​a user, a terminal (smart speaker), and a server. The user speaks to the terminal, the terminal sends voice data to the server, and the server analyzes and processes the voice data to generate an appropriate response.

[0867] When a user speaks into a smart speaker, the device uses its microphone to capture the user's voice. Specifically, the user might say things like "Good morning" or "What did you do today?". This voice data is captured by the device and sent to a server via the internet.

[0868] The transmitted audio data is first temporarily stored on the server. The server converts the stored audio data into text using speech recognition technology. Services such as the Google Speech-to-Text API are used for this process. The text data obtained through speech recognition technology is stored in a database on the server.

[0869] Next, the server extracts speech features (e.g., pronunciation patterns, intonation) from the audio data. These speech features are used as data to train a deep learning model (e.g., TensorFlow). During the training process, the server learns the user's unique speaking style and intonation, and builds a model for generating natural-sounding dialogue.

[0870] When the user speaks a question or has a conversation into the device again, the device records the audio and sends it to the server. The server converts the audio back into text and performs morphological analysis (e.g., using MeCab). Based on the analyzed text data, the server uses a trained deep learning model to generate an appropriate response. For example, it might generate a response such as, "I had a relaxing day."

[0871] The generated text responses are converted into audio data using speech synthesis technology. Speech synthesis services such as the Google Text-to-Speech API are used for this process. The server then generates lip-sync data synchronized with the audio data to recreate the avatar's mouth movements.

[0872] The device uses the received audio and lip-sync data to respond to the user. The user can listen to the avatar's response, such as "I had a relaxing day," while observing the avatar's natural mouth movements.

[0873] Specific example

[0874] One day, a user asks a smart speaker, "What's the weather like tomorrow?" The device records the voice and sends it to a server. The server converts the voice to text, performs further analysis, and generates an appropriate response. This response is "It will be sunny tomorrow," and is converted into voice data via speech synthesis technology. The generated voice data and lip-sync data are sent to the device, and an avatar speaks "It will be sunny tomorrow," with its mouth movements naturally reproduced.

[0875] Example of a prompt

[0876] 1. "What's the weather like tomorrow?"

[0877] 2. "Please tell me about the latest news."

[0878] 3. "What fun things happened to you today?"

[0879] This system generates responses where voice and mouth movements are naturally synchronized, providing users with a sense of psychological reassurance and familiarity.

[0880] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0881] Step 1:

[0882] The user speaks to the smart speaker.

[0883] Input: Voice (Example: "Good morning")

[0884] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0885] Output: Recorded audio data

[0886] Step 2:

[0887] The device (smart speaker) sends the recorded audio data to the server.

[0888] Input: Recorded audio data

[0889] Specific operation: Captured audio data is sent to the server via the internet.

[0890] Output: Audio data sent to the server

[0891] Step 3:

[0892] The server temporarily stores the received audio data.

[0893] Input: Audio data sent to the server

[0894] Specific operation: The audio data is saved to storage (e.g., Amazon S3).

[0895] Output: Saved audio data

[0896] Step 4:

[0897] The server uses speech recognition technology to convert the audio data into text.

[0898] Input: Saved audio data

[0899] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and text data is returned.

[0900] Output: Text data

[0901] Step 5:

[0902] The server extracts speech features from the audio data and stores them in a database.

[0903] Input: Saved audio data

[0904] Specific operation: Audio data is analyzed, features such as pronunciation patterns and intonation are extracted, and these are stored in a database.

[0905] Output: Speech feature data

[0906] Step 6:

[0907] The server uses speech features to train a deep learning model.

[0908] Input: Speech feature data

[0909] Specific operation: A neural network is trained using a framework such as TensorFlow with audio feature data.

[0910] Output: Trained deep learning model

[0911] Step 7:

[0912] The user speaks to the smart speaker, asking questions or engaging in conversation.

[0913] Input: Voice (Example: "What did you do today?")

[0914] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[0915] Output: Recorded audio data

[0916] Step 8:

[0917] The device sends the recorded audio data to the server.

[0918] Input: Recorded audio data

[0919] Specific operation: Captured audio data is sent to the server via the internet.

[0920] Output: Audio data sent to the server

[0921] Step 9:

[0922] The server converts the audio data into text and performs morphological analysis.

[0923] Input: Audio data sent to the server

[0924] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and after obtaining text data, morphological analysis is performed using MeCab, etc.

[0925] Output: Morphologically analyzed text data

[0926] Step 10:

[0927] The server uses a trained deep learning model to generate appropriate responses to questions and conversations.

[0928] Input: Morphologically analyzed text data

[0929] Specific operation: Input text data into a trained deep learning model and generate an appropriate response.

[0930] Output: Generated text response

[0931] Step 11:

[0932] The server converts the generated text response into speech data using speech synthesis technology.

[0933] Input: Generated text response

[0934] Specific operation: Convert text responses into speech data using the Google Text-to-Speech API, etc.

[0935] Output: Generated audio data

[0936] Step 12:

[0937] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[0938] Input: Generated audio data

[0939] Specific operation: Analyzes audio data and creates lip-sync data to match the audio.

[0940] Output: Lip-sync data

[0941] Step 13:

[0942] The device receives voice data and lip-sync data and displays a response to the user.

[0943] Input: Generated audio data and lip-sync data

[0944] Specific operation: Based on the received data, the avatar responds to the user with voice and simultaneously displays the movement of its mouth.

[0945] Output: Voice responses and mouth movements by the avatar

[0946] In this way, through each processing step, users can enjoy natural conversations with specific individuals in remote locations through the reproduction of natural dialogue and the accompanying mouth movements of their avatars.

[0947] (Application Example 1)

[0948] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0949] Conventional speech recognition and voice response systems have limitations in achieving natural dialogue with users, and systems combining speech synthesis and lip-syncing have not been widely implemented. Furthermore, in food delivery services, it has been difficult to provide personalized menu suggestions and answer questions through dialogue with users. This invention aims to solve these problems and enable users to use food delivery services more comfortably.

[0950] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0951] In this invention, the server includes means for collecting voice samples from a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, virtual assistant means for analyzing the user's questions, generating an appropriate response, and outputting that response using speech synthesis technology, and means for using a generative AI model that generates a response based on the analysis results of the user's voice input. This enables a more natural and intuitive conversation when a user uses a food delivery service.

[0952] A "voice sample" is data that is a recording of a person speaking in a remote location.

[0953] "Speech features" refer to characteristic information such as pronunciation patterns and intonation extracted from collected speech samples.

[0954] A "deep learning model" is a model that uses a multi-layered neural network to learn from large amounts of data and perform predictions and classifications on new data.

[0955] A "virtual assistant" is a software system that generates appropriate responses to user voice input and outputs those responses using speech synthesis technology.

[0956] A "generative AI model" is an artificial intelligence model trained to generate responses based on a user's voice input.

[0957] A "lip-syncing method" is a means of displaying the mouth movements of an avatar in sync with the generated audio.

[0958] A "food delivery service" is a service that delivers food to customers based on their orders.

[0959] This invention is a system that collects voice samples from people in remote locations and uses them to provide users with natural, real-time conversations. This system is particularly effective in food delivery services, where users can use devices such as smartphones or smart glasses to interact with a virtual assistant and obtain information about menus and recommended dishes.

[0960] First, the server collects voice samples and temporarily stores the data. These voice samples are obtained when the user speaks into their smartphone or smart glasses. For example, if the user asks, "What pasta dish do you recommend?", the voice is recorded by the device and sent to the server.

[0961] Next, the server analyzes the collected audio data and converts it into text using speech recognition technology. Simultaneously, it extracts audio features (such as pronunciation patterns and intonation) and stores them in a database. This allows a deep learning model to be trained using the audio features.

[0962] A trained deep learning model generates appropriate responses to voice input from the user. When the user provides voice input again, the server analyzes the speech to understand the intent and uses the generative AI model to generate an appropriate response. This response is then converted into audio data using speech synthesis technology and output to the user.

[0963] Furthermore, lip-sync data synchronized with this generated audio data is created, allowing the virtual assistant avatar to reproduce natural mouth movements. This allows users to visually confirm that the virtual assistant is speaking naturally.

[0964] As a concrete example, when a user asks, "What pasta dish do you recommend?", the system generates a response through the following process: It replies with voice, "My recommendation is carbonara," and an avatar makes the same response with natural mouth movements. This allows the user to experience an immersive conversation.

[0965] The hardware used will be a smartphone, smart glasses, and a server. The software used will be speech recognition technology (e.g., the speech_recognition library), speech synthesis technology (e.g., the gtts library), and a generative AI model (e.g., the GPT-2 model from the transformers library).

[0966] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0967] Step 1:

[0968] When a user speaks to their smartphone or smart glasses, their voice is recorded by the device. The recorded voice data is then sent directly to the server. For example, if a user asks, "What pasta dish do you recommend?", that voice will be collected. The input is the user's voice data, and the output is the voice data sent to the server.

[0969] Step 2:

[0970] The server temporarily stores the received audio data. Then, it converts the audio data into text data using speech recognition technology. Here, the `speech_recognition` library is used to convert speech to text. The input is audio data, and the output is text data.

[0971] Step 3:

[0972] The server further analyzes the text data and extracts speech features (pronunciation patterns, intonation, etc.). A speech feature extraction algorithm is used for the analysis, and the extracted features are stored in a database. The input is text data, and the output is speech feature data.

[0973] Step 4:

[0974] The server accumulates speech feature data and uses it to train a deep learning model. A large amount of speech features are used for training, and a trained generative AI model is constructed. The input is speech feature data, and the output is the trained generative AI model.

[0975] Step 5:

[0976] When the user speaks a question to the device again, the device records the audio and sends it to the server. The server analyzes the received audio, converts it into text data, and performs morphological analysis to understand the user's intent. The input is the newly collected audio data, and the output is the analyzed text data and the results of the intent analysis.

[0977] Step 6:

[0978] The server uses a trained generative AI model to generate appropriate responses to user voice input. The generated responses are converted into speech data using speech synthesis technology (e.g., the GTTS library). The input is the intent analysis result, and the output is the speech response data.

[0979] Step 7:

[0980] The server generates lip-sync data synchronized with the generated audio data, ensuring that the virtual assistant avatar reflects natural mouth movements. The lip-sync data is used to control the avatar's mouth movements in real time. The input is audio data, and the output is lip-sync data.

[0981] Step 8:

[0982] The terminal receives audio data and lip-sync data from the server, plays an audio response to the user, and displays the avatar's lip-sync. The input is audio data and lip-sync data, and the output is the audio response to the user and the avatar's natural mouth movements.

[0983] Through the steps outlined above, users can ask questions about the food delivery service in a natural, conversational format. This allows users to obtain detailed information through dialogue, significantly improving convenience.

[0984] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0985] This invention is a system that reproduces the voice characteristics of a specific person in a remote location, recognizes the user's emotions, and adjusts the response accordingly. The system consists of a smart speaker and a server, and each process is performed as follows.

[0986] 1. Collection and storage of audio samples:

[0987] The user speaks to the smart speaker, which provides a voice sample of a specific person. For example, the user might say "Good morning."

[0988] The device (smart speaker) records the user's voice and sends that voice data to the server.

[0989] 2. Analysis of audio data:

[0990] The server stores the received audio data and converts it into text data using speech recognition technology. Based on the converted text, it extracts audio features such as pronunciation patterns and intonation.

[0991] The extracted speech features are stored in a database and used to train a deep learning model.

[0992] 3. Training deep learning models:

[0993] The server uses the accumulated speech features to train a deep learning model to learn specific speaking styles and intonations.

[0994] The trained model will be used to generate future user responses.

[0995] 4. User emotion recognition:

[0996] When a user speaks a new question or engages in a conversation with the smart speaker, the device records the audio again and sends it to the server.

[0997] The server analyzes the received audio data and uses emotion recognition technology to analyze the user's emotions.

[0998] 5. Analyzing the intent behind questions and conversations:

[0999] The server converts the audio data into text, performs morphological analysis, and understands the user's intent.

[1000] 6. Response generation and adjustment:

[1001] Using a pre-trained deep learning model, it generates appropriate responses to user questions. For example, it can generate a response like, "I had a relaxing day today."

[1002] The response is adjusted based on the user's emotions as assessed by the emotion recognition engine. For example, if the user appears sad, it might generate a response such as, "I had a relaxing day today. You seem a little down, are you okay?"

[1003] The generated text responses are converted into speech data using text-to-speech (TTS) technology.

[1004] 7. Lip sync and avatar display:

[1005] The server generates lip-sync data corresponding to the audio data, naturally expressing the avatar's mouth movements and facial expressions.

[1006] The device receives audio data and lip-sync data, plays the audio, and displays the avatar's facial expressions and mouth movements.

[1007] 8. User experience:

[1008] Users can listen to the responses generated through the smart speaker and see the avatar's natural mouth movements and facial expressions.

[1009] This allows users to enjoy natural conversations with family members or deceased loved ones who live far away, and to receive responses that are sensitive to the user's emotions.

[1010] In this way, the present invention realizes a system that realistically reproduces the voice characteristics of a specific person, understands the user's emotions, and provides a more natural and approachable conversational experience.

[1011] The following describes the processing flow.

[1012] Step 1:

[1013] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[1014] Step 2:

[1015] The device (smart speaker) records the user's voice and saves it to a temporary buffer. It then sends the recorded voice data to the server.

[1016] Step 3:

[1017] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[1018] Step 4:

[1019] The server converts the audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features (such as pronunciation patterns and intonation).

[1020] Step 5:

[1021] The server stores the extracted speech features in a database. Simultaneously, this data is used to train a deep learning model.

[1022] Step 6:

[1023] The user initiates a new question or conversation with the smart speaker. For example, they might ask, "How was your day?"

[1024] Step 7:

[1025] The device (smart speaker) re-records the user's voice and sends it to the server.

[1026] Step 8:

[1027] The server analyzes the received audio data and converts it back into text data. This text is then subjected to morphological analysis to determine the user's intent.

[1028] Step 9:

[1029] The server uses an emotion recognition engine during the process of analyzing voice data to analyze the user's emotions (such as joy, anger, sadness, etc.).

[1030] Step 10:

[1031] The server evaluates the user's emotions and uses a trained deep learning model to generate an appropriate response. For example, it might generate a response like, "I had a relaxing day."

[1032] Step 11:

[1033] The server adjusts its response based on the emotion recognition results. For example, if the user seems sad, it might change the response to something like, "I had a relaxing day today. You seem a little down, are you okay?"

[1034] Step 12:

[1035] The generated text response is converted into speech data using text-to-speech (TTS) technology. This makes the response sound more natural.

[1036] Step 13:

[1037] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's mouth movements and facial expressions.

[1038] Step 14:

[1039] The server sends audio data and lip-sync data to the terminal.

[1040] Step 15:

[1041] The device (smart speaker) receives audio data and lip-sync data from the server. It plays the audio and displays the avatar's facial expressions and mouth movements.

[1042] Step 16:

[1043] Users listen to the generated responses and observe the avatar's natural mouth movements and facial expressions. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[1044] This specific processing flow allows the present invention to faithfully reproduce the voice characteristics and pronunciation patterns of a particular person and provide a friendly response that corresponds to the user's emotions.

[1045] (Example 2)

[1046] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1047] Providing a natural and engaging conversational experience with a specific person located remotely is challenging. Furthermore, recognizing the user's emotions and generating appropriate responses accordingly is also difficult. Therefore, the challenge lies in accurately reproducing a specific person's voice and speaking style while simultaneously providing responses that resonate with the user's emotions.

[1048] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1049] In this invention, the server includes means for collecting voice samples of a specific person located remotely, means for analyzing the collected voice samples to extract voice features such as pronunciation patterns and intonation, means for training a deep learning model using the voice features, means for recognizing the user's emotions and adjusting the response content, and means for converting voice input from the user into text and analyzing the intent. This makes it possible to realistically reproduce the voice of a specific person and provide a natural conversational experience that responds to the user's emotions.

[1050] A "voice sample" is a digital recording of a user's speech, which imitates the voice of a specific person.

[1051] "Pronunciation patterns" refer to characteristics of speech such as the flow of phonemes, rhythm, stress, and accentuation.

[1052] "Intonation" refers to the changes in pitch and intonation of a voice in spoken language.

[1053] "Speech features" are specific data points extracted from speech samples, and include elements such as pronunciation patterns, intonation, pitch, and length.

[1054] A "deep learning model" is a type of artificial intelligence trained using large amounts of data, and specifically refers to models that use neural networks.

[1055] "Emotion recognition" refers to technology that analyzes and identifies a user's emotional state from their voice or text.

[1056] An "avatar" refers to a representation of a person or character displayed in a virtual space for interaction with a user.

[1057] "Lip sync" refers to a technology that synchronizes the mouth movements of an avatar with the audio.

[1058] "Speech recognition technology" is a technology for converting speech into text, and includes the process of analyzing a speech sample and converting it into text data.

[1059] "Speech synthesis technology" refers to the technology that generates speech based on text data.

[1060] Modes for carrying out the invention

[1061] This invention relates to a system that reproduces the voice characteristics of a specific person located remotely, recognizes the user's emotions, and adjusts its response accordingly. The system consists of a smart speaker and a server and includes the following elements:

[1062] 1. Collection and storage of audio samples

[1063] Voice samples are collected when the user speaks to the smart speaker. For example, the user says "Good morning." The device (smart speaker) records the user's voice with its microphone and temporarily stores this voice data in its internal memory. Next, the device sends this voice data to the server via Wi-Fi.

[1064] 2. Analysis of audio data

[1065] The server stores the received audio data in a dedicated database. Then, it uses speech recognition technology to convert the audio data into text data. Specifically, it uses the Google Speech-to-Text API to convert the audio to text. Based on this text data, it extracts speech features such as pronunciation patterns and intonation, and stores them in the database.

[1066] 3. Training of deep learning models

[1067] A deep learning model is trained using speech features. The server uses a framework such as TensorFlow to learn specific speech patterns and intonations. This training requires a large amount of speech data and corresponding text data. Once training is complete, the model is saved to a specific folder.

[1068] 4. User emotion recognition

[1069] When the user speaks a question or engages in conversation with the smart speaker again, the device records the voice again and sends the audio data to the server. The server analyzes the received audio data and uses emotion recognition technology (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. This analysis result is stored in JSON format.

[1070] 5. Analyzing the intent behind questions and conversations

[1071] The server converts the audio data back into text and performs morphological analysis. For example, MeCab is used for this purpose. Morphological analysis helps understand the user's intent and obtain the information necessary to generate an appropriate response.

[1072] 6. Response generation and adjustment

[1073] Using a pre-trained deep learning model, the server generates appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, considering the user's mood, it might generate additional follow-up responses such as, "You seem a little down, are you okay?" The generated text responses are then converted into speech data using the Google Text-to-Speech API.

[1074] 7. Lip sync and avatar display

[1075] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's facial expressions and mouth movements. The terminal receives this audio data and lip-sync data, playing the audio while simultaneously displaying the avatar's facial expressions and mouth movements on the screen.

[1076] 8. User experience

[1077] Users can hear voice responses from smart speakers and see the natural facial expressions and mouth movements of the avatar on the display. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[1078] Specific example

[1079] For example, if a user says "Good morning" to a smart speaker, the smart speaker records the voice and sends it to a server. The server analyzes this voice, extracts features, and stores them in a database. Then, it trains a deep learning model using TensorFlow. If the user asks another question, such as "How was your day?", the server uses emotion recognition technology to analyze the user's emotions. For example, if the user sounds sad, it might generate a response like, "I had a relaxing day. You seem a little down, are you okay?" and convert it into speech using the Google Text-to-Speech API. The device plays this speech, and the display shows the avatar's facial expressions and mouth movements.

[1080] Example of a prompt

[1081] Please generate example responses to the question, "How was your day?" Include additional, emotionally sensitive follow-up if the user seems sad.

[1082] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1083] Step 1: Collect and save audio samples

[1084] The user imitates a specific person's voice and says "Good morning" to the smart speaker. The smart speaker, as the device, records this voice using its microphone. The recorded voice data is temporarily stored in internal memory and then sent to a server via Wi-Fi.

[1085] Input: User's voice "Good morning"

[1086] Output: Recorded audio data (digital format)

[1087] Step 2: Analysis of the voice moon

[1088] The server stores the received audio data and uses speech recognition technology to convert the audio into text data. Using the Google Speech-to-Text API, the audio data is converted into the text "Good morning." Based on this text data, speech features such as pronunciation patterns and intonation are extracted and stored in a new database.

[1089] Input: Recorded audio data

[1090] Output: Text data and speech features

[1091] Step 3: Training the deep learning model

[1092] The server trains a deep learning model using speech features. Libraries such as TensorFlow are used to learn specific speaking styles and intonations. In this training example, a large amount of speech data and corresponding text data are used to enable the model to mimic specific speaking styles and intonations. The trained model is saved to a specific folder.

[1093] Input: Speech features

[1094] Output: Trained deep learning model

[1095] Step 4: User emotion recognition

[1096] The user speaks to the smart speaker again, for example, asking, "How was your day?" The device records this voice and sends it to the server. The server receives the voice data and analyzes the user's emotions using emotion recognition technology (e.g., IBM Watson Tone Analyzer). The analysis results are stored in JSON format.

[1097] Input: User's voice "How was your day?"

[1098] Output: User sentiment analysis results

[1099] Step 5: Analyzing the intent behind questions and conversations

[1100] The server then converts the received audio data back into text data and performs morphological analysis. This is done using a morphological analyzer such as MeCab. By converting the audio into text data such as "How was your day?" and performing morphological analysis, the server understands what the user is asking.

[1101] Input: Audio data

[1102] Output: Analyzed text data and user intent

[1103] Step 6: Generating and adjusting the response

[1104] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, based on the sentiment recognition results, it might generate a follow-up response such as, "You seem a little down, are you okay?" This generated response is then converted to speech using the Google Text-to-Speech API.

[1105] Input: Analyzed text data and user sentiment analysis results

[1106] Output: Adjusted voice response

[1107] Step 7: Lip sync and avatar display

[1108] The server generates lip-sync data that synchronizes with the generated audio data. This lip-sync data naturally expresses the avatar's mouth movements and facial expressions. The terminal receives the audio data and lip-sync data, plays the audio, and simultaneously displays the avatar on the screen. The avatar's mouth moves in accordance with the lip-sync data, displaying appropriate facial expressions.

[1109] Input: Adjusted voice response

[1110] Output: Avatar facial expressions corresponding to audio and lip-sync data

[1111] Step 8: User Experience

[1112] Users can hear voice responses played from the smart speaker. They can also feel like they are actually having a conversation by observing the natural mouth movements and facial expressions of the avatar displayed on the screen. This system is designed to allow users to enjoy a natural conversational experience with a specific person.

[1113] Input: User's question

[1114] Output: Natural conversation experience and avatar display

[1115] (Application Example 2)

[1116] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1117] A system that can reproduce the voice of a specific person in a remote location and adjust its response by recognizing the user's emotions can provide a natural, emotionally resonant conversational experience in the home and in daily life. However, while such a system could also be applied to work environments such as factories to provide emotional support to employees and improve work efficiency, current systems do not adequately address this need. There is a need for a new system that allows factory workers to relax and boost morale through direct interaction.

[1118] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice samples of a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, and means for recognizing the user's emotions and adjusting the response based on the results. This makes it possible to provide a natural conversational experience that is empathetic to emotions, even in work environments such as factories, enabling mental support for employees and improvement of work efficiency.

[1119] A "voice sample" refers to audio data recorded for use in other systems or algorithms, specifically recording the voice of a particular person.

[1120] "Speech features" refer to characteristic data representations extracted from speech data, such as pronunciation patterns and intonation.

[1121] A "deep learning model" refers to a type of machine learning model that is constructed using multi-layer neural networks to learn complex patterns from data.

[1122] "User voice input" refers to the act of a system user entering information by voice, or the input data itself.

[1123] "Response" refers to the reply or content of the response that the system generates in response to voice input from the user.

[1124] An "avatar" refers to a virtual person or character displayed on a computer screen or by a robot, whose mouth movements and facial expressions are synchronized with the voice.

[1125] "Emotion recognition" refers to a technology or process that analyzes audio data and other information to identify the emotional state of a speaker.

[1126] A "server" refers to a computer system that stores and processes data on a network and communicates with client devices.

[1127] "Means of collection" refers to hardware and software for recording or collecting audio samples via a network.

[1128] "Means of analysis" refers to algorithms and software used to analyze collected data and extract and process speech features.

[1129] "Training methods" refer to the process and techniques of optimizing deep learning models using speech features to improve their skills so that they can accurately generate responses.

[1130] "Means of output" refers to speakers or speech synthesis technology that allow the user to hear the generated response as audio.

[1131] "Means of display" refers to monitors and display technologies used to show users the mouth movements and facial expressions of avatars.

[1132] "Means of adjustment" refers to algorithms and technologies that adjust responses generated based on the results of emotion recognition to create appropriate responses that match the user's emotions.

[1133] This invention provides a system for use in work environments such as factories that reproduces the voice of a specific person, recognizes the emotions of employees, and generates appropriate responses. This can provide emotional support to employees and improve work efficiency.

[1134] The server stores collected audio samples and converts them into text data using speech recognition technology. It also extracts speech features such as pronunciation patterns and intonation, stores them in a database, and uses them to train a deep learning model. The server trains the deep learning model to create a model that learns specific speaking styles and intonations. Then, it uses the trained model to generate appropriate responses to user questions and outputs them as audio. For audio output, text-to-speech (TTS) technology is used to convert the text responses into audio data.

[1135] Furthermore, the server recognizes the user's emotions and adjusts the response generated based on that. For example, if the user appears sad, it generates a response that takes those emotions into consideration. In this way, a natural conversational experience is achieved.

[1136] The terminal (such as a robot or monitor) uses audio and lip-sync data received from the server to play the audio and display the avatar's facial expressions and mouth movements. This allows the user to hear the audio response and see the avatar's natural mouth movements and expressions.

[1137] Specifically, the `speech_recognition` library is used for speech recognition, and the `transformers` library is used for emotion recognition. The `pyttsx3` library is used to convert text to speech.

[1138] The following are specific examples of prompt statements.

[1139] "Please generate an appropriate response based on my emotions. Analyze the following text to recognize my emotions and then provide a friendly voice response accordingly."

[1140] In this way, the present invention provides a natural conversational experience that is attentive to the emotions of employees, even in work environments such as factories, thereby achieving emotional support and improved work efficiency.

[1141] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1142] Step 1:

[1143] The user provides a voice sample. The user speaks to the robot and provides a voice sample of a specific person. For example, the user might say "Good morning." This voice sample is recorded by the device and sent to the server.

[1144] Step 2:

[1145] The server analyzes the audio sample. The server stores the received audio data and converts it into text data using speech recognition technology (e.g., the speech_recognition library). Based on the converted text, it extracts speech features such as pronunciation patterns and intonation. These features are stored in a database.

[1146] Step 3:

[1147] The server trains a deep learning model. The server uses accumulated speech features to train a deep learning model to learn specific speech patterns and intonations. During this process, the transformers library is used to optimize the model using speech features as input. The trained model is then used to generate future user responses.

[1148] Step 4:

[1149] The system recognizes the user's emotions. When the user speaks to the robot again, the terminal records the voice again and sends it to the server. The server analyzes the received voice data and uses emotion recognition technology (e.g., an emotion recognition model from the transformers library) to analyze the user's emotions. The emotion label is output from the voice data as input.

[1150] Step 5:

[1151] The system analyzes the intent behind questions and conversations. The server converts the audio data into text (for example, using the speech_recognition library) and performs morphological analysis to understand the user's intent. This analysis outputs the intent in text format, and a response is determined based on it.

[1152] Step 6:

[1153] The server generates a response and outputs it as speech. A trained deep learning model is used to generate an appropriate response to the user's question. For example, a response such as "I had a relaxing day today" might be generated. The generated text response is converted into speech data using speech synthesis technology (e.g., the pyttsx3 library) and output.

[1154] Step 7:

[1155] The device performs lip-syncing and displays the avatar. The server generates lip-sync data corresponding to the audio data (for example, using an animation library) to create data that naturally expresses the avatar's mouth movements and facial expressions. The device uses this data to play the audio while displaying the avatar's mouth movements and facial expressions.

[1156] Step 8:

[1157] The user experience is enhanced. Users can listen to responses generated through the robot and see the avatar's natural mouth movements and facial expressions. This allows users to enjoy a natural conversational experience with family members or deceased loved ones who are far away, and receive responses that are more emotionally resonant.

[1158] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1159] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1160] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1161] [Fourth Embodiment]

[1162] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1163] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1164] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1165] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1166] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1167] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1168] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1169] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1170] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1171] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1172] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1173] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1174] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1175] The present invention mainly involves the following processing steps.

[1176] 1. Collection of audio samples:

[1177] Voice samples are collected when users speak to a smart speaker.

[1178] The device (smart speaker) records the user's voice and sends that data to the server.

[1179] For example, when a user says "Good morning," the voice is recorded by the smart speaker, and the voice data is sent to the server.

[1180] 2. Storage and analysis of audio data:

[1181] The server temporarily stores the received audio data.

[1182] The saved audio data is converted into text using speech recognition technology. Additionally, audio features (pronunciation patterns, intonation, etc.) are extracted.

[1183] The extracted speech features are stored in a database and used later for deep learning.

[1184] 3. Training deep learning models:

[1185] The server uses the accumulated audio features to train a deep learning model.

[1186] During the training process, specific speaking styles and intonations are learned, and models are built for generating natural dialogue.

[1187] For example, phrases that users repeatedly say, such as "Hello" or "How are you?", can be used as training data.

[1188] 4. Receiving and analyzing questions and conversations:

[1189] When a user speaks a question or engages in a conversation with a smart speaker, the device records it and sends it to a server.

[1190] The server analyzes the received audio, converts it to text, and performs morphological analysis to understand the user's intent.

[1191] For example, if a user asks "What did you do today?", the content of that question is converted into text and analyzed.

[1192] 5. Response generation and speech synthesis:

[1193] A trained deep learning model is used to generate appropriate responses to user questions.

[1194] The generated response is converted into speech data using text-to-speech (TTS) technology.

[1195] For example, a response like "I had a relaxing day today" might be generated.

[1196] 6. Lip sync and avatar display:

[1197] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[1198] The terminal receives voice data and lip-sync data, and uses them to actually respond to the user.

[1199] The user can hear the response and simultaneously see the avatar's natural mouth movements.

[1200] For example, an avatar saying "I had a relaxing day today" will be displayed with accurate mouth movements.

[1201] Thus, the present invention provides a system that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away, by combining a technology that reproduces the voice characteristics of a specific person with a lip-sync technology that synchronizes with the voice. It also includes specific means to make the user feel psychologically secure and familiar with the system.

[1202] The following describes the processing flow.

[1203] Step 1:

[1204] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[1205] Step 2:

[1206] The device (smart speaker) records the user's voice and saves it to a temporary buffer. After recording is complete, it sends the voice data to the server.

[1207] Step 3:

[1208] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[1209] Step 4:

[1210] The server converts audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features, including pronunciation patterns and intonation.

[1211] Step 5:

[1212] The server stores the extracted audio features in a database. Simultaneously, it uses this data to train a deep learning model. A multi-layered neural network is used for training.

[1213] Step 6:

[1214] The user initiates a new conversation or asks a question to the smart speaker. For example, they might ask, "How was your day?"

[1215] Step 7:

[1216] The device (smart speaker) re-records the user's voice and sends it to the server.

[1217] Step 8:

[1218] The server receives the audio data and converts it into text data. It then performs morphological analysis on the text to analyze the user's intent.

[1219] Step 9:

[1220] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a response like, "I had a relaxing day."

[1221] Step 10:

[1222] The server converts the generated text response into speech data using text-to-speech (TTS) technology. This enables natural-sounding speech.

[1223] Step 11:

[1224] The server generates lip-sync data corresponding to the audio data. This allows the audio and the avatar's mouth movements to be synchronized.

[1225] Step 12:

[1226] The server sends audio data and lip-sync data to the terminal.

[1227] Step 13:

[1228] The device (smart speaker) receives audio data and lip-sync data from the server. Based on the received data, it plays the audio and simultaneously displays the mouth movements of the avatar.

[1229] Step 14:

[1230] Users listen to responses generated through smart speakers and observe the natural mouth movements of their avatars. This allows them to enjoy a natural conversational experience with loved ones.

[1231] In this way, a system is created that allows users to enjoy natural conversations with family members or deceased loved ones who are located far away.

[1232] (Example 1)

[1233] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1234] Conventional speech recognition and dialogue generation systems struggle to reproduce the natural speaking style and intonation of a specific person located remotely. As a result, they are insufficient as systems that provide users with a sense of psychological security and familiarity. Furthermore, these systems often lack visual feedback for the generated speech, leading to mismatches between speech and mouth movements. This, in turn, negatively impacts the user experience.

[1235] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1236] In this invention, the server includes means for acquiring voice data, means for storing the acquired voice data, means for converting the stored voice data into text using speech recognition technology, means for extracting and accumulating voice features, means for training a deep learning model using the accumulated voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice using speech synthesis technology, and means for displaying the mouth movements of an avatar in synchronization with the output voice. This makes it possible to generate a response in which the voice and mouth movements are naturally synchronized, providing the user with a sense of psychological security and familiarity.

[1237] "Means of acquisition" refers to devices or mechanisms for collecting the user's voice.

[1238] "Means of storage" refers to a storage device or service for temporarily storing the collected audio data.

[1239] "Speech recognition technology" refers to algorithms and software used to convert speech data into text data.

[1240] "Speech features" are identifiable characteristics such as pronunciation patterns and intonation that are extracted from speech data.

[1241] "Means of storage" refer to databases or storage devices for saving and managing extracted speech features.

[1242] A "deep learning model" is a neural network model that is trained using speech features.

[1243] "Means for generating responses in response to voice input" refers to a mechanism that uses a trained deep learning model to create appropriate responses to user voice input.

[1244] "Speech synthesis technology" refers to algorithms and software used to convert generated text data into speech data.

[1245] "Means for displaying the mouth movements of an avatar" refers to a display device or software that reproduces the mouth movements of an avatar in synchronization with the outputted audio.

[1246] To implement this invention, three elements are necessary: ​​a user, a terminal (smart speaker), and a server. The user speaks to the terminal, the terminal sends voice data to the server, and the server analyzes and processes the voice data to generate an appropriate response.

[1247] When a user speaks into a smart speaker, the device uses its microphone to capture the user's voice. Specifically, the user might say things like "Good morning" or "What did you do today?". This voice data is captured by the device and sent to a server via the internet.

[1248] The transmitted audio data is first temporarily stored on the server. The server converts the stored audio data into text using speech recognition technology. Services such as the Google Speech-to-Text API are used for this process. The text data obtained through speech recognition technology is stored in a database on the server.

[1249] Next, the server extracts speech features (e.g., pronunciation patterns, intonation) from the audio data. These speech features are used as data to train a deep learning model (e.g., TensorFlow). During the training process, the server learns the user's unique speaking style and intonation, and builds a model for generating natural-sounding dialogue.

[1250] When the user speaks a question or has a conversation into the device again, the device records the audio and sends it to the server. The server converts the audio back into text and performs morphological analysis (e.g., using MeCab). Based on the analyzed text data, the server uses a trained deep learning model to generate an appropriate response. For example, it might generate a response such as, "I had a relaxing day."

[1251] The generated text responses are converted into audio data using speech synthesis technology. Speech synthesis services such as the Google Text-to-Speech API are used for this process. The server then generates lip-sync data synchronized with the audio data to recreate the avatar's mouth movements.

[1252] The device uses the received audio and lip-sync data to respond to the user. The user can listen to the avatar's response, such as "I had a relaxing day," while observing the avatar's natural mouth movements.

[1253] Specific example

[1254] One day, a user asks a smart speaker, "What's the weather like tomorrow?" The device records the voice and sends it to a server. The server converts the voice to text, performs further analysis, and generates an appropriate response. This response is "It will be sunny tomorrow," and is converted into voice data via speech synthesis technology. The generated voice data and lip-sync data are sent to the device, and an avatar speaks "It will be sunny tomorrow," with its mouth movements naturally reproduced.

[1255] Example of a prompt

[1256] 1. "What's the weather like tomorrow?"

[1257] 2. "Please tell me about the latest news."

[1258] 3. "What fun things happened to you today?"

[1259] This system generates responses where voice and mouth movements are naturally synchronized, providing users with a sense of psychological reassurance and familiarity.

[1260] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1261] Step 1:

[1262] The user speaks to the smart speaker.

[1263] Input: Voice (Example: "Good morning")

[1264] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[1265] Output: Recorded audio data

[1266] Step 2:

[1267] The device (smart speaker) sends the recorded audio data to the server.

[1268] Input: Recorded audio data

[1269] Specific operation: Captured audio data is sent to the server via the internet.

[1270] Output: Audio data sent to the server

[1271] Step 3:

[1272] The server temporarily stores the received audio data.

[1273] Input: Audio data sent to the server

[1274] Specific operation: The audio data is saved to storage (e.g., Amazon S3).

[1275] Output: Saved audio data

[1276] Step 4:

[1277] The server uses speech recognition technology to convert the audio data into text.

[1278] Input: Saved audio data

[1279] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and text data is returned.

[1280] Output: Text data

[1281] Step 5:

[1282] The server extracts speech features from the audio data and stores them in a database.

[1283] Input: Saved audio data

[1284] Specific operation: Audio data is analyzed, features such as pronunciation patterns and intonation are extracted, and these are stored in a database.

[1285] Output: Speech feature data

[1286] Step 6:

[1287] The server uses speech features to train a deep learning model.

[1288] Input: Speech feature data

[1289] Specific operation: A neural network is trained using a framework such as TensorFlow with audio feature data.

[1290] Output: Trained deep learning model

[1291] Step 7:

[1292] The user speaks to the smart speaker, asking questions or engaging in conversation.

[1293] Input: Voice (Example: "What did you do today?")

[1294] Specific operation: The user speaks, and the smart speaker captures it with its microphone.

[1295] Output: Recorded audio data

[1296] Step 8:

[1297] The device sends the recorded audio data to the server.

[1298] Input: Recorded audio data

[1299] Specific operation: Captured audio data is sent to the server via the internet.

[1300] Output: Audio data sent to the server

[1301] Step 9:

[1302] The server converts the audio data into text and performs morphological analysis.

[1303] Input: Audio data sent to the server

[1304] Specific operation: Audio data is sent to the Google Speech-to-Text API, etc., and after obtaining text data, morphological analysis is performed using MeCab, etc.

[1305] Output: Morphologically analyzed text data

[1306] Step 10:

[1307] The server uses a trained deep learning model to generate appropriate responses to questions and conversations.

[1308] Input: Morphologically analyzed text data

[1309] Specific operation: Input text data into a trained deep learning model and generate an appropriate response.

[1310] Output: Generated text response

[1311] Step 11:

[1312] The server converts the generated text response into speech data using speech synthesis technology.

[1313] Input: Generated text response

[1314] Specific operation: Convert text responses into speech data using the Google Text-to-Speech API, etc.

[1315] Output: Generated audio data

[1316] Step 12:

[1317] The server generates lip-sync data synchronized with the audio data to reproduce the avatar's mouth movements.

[1318] Input: Generated audio data

[1319] Specific operation: Analyzes audio data and creates lip-sync data to match the audio.

[1320] Output: Lip-sync data

[1321] Step 13:

[1322] The device receives voice data and lip-sync data and displays a response to the user.

[1323] Input: Generated audio data and lip-sync data

[1324] Specific operation: Based on the received data, the avatar responds to the user with voice and simultaneously displays the movement of its mouth.

[1325] Output: Voice responses and mouth movements by the avatar

[1326] In this way, through each processing step, users can enjoy natural conversations with specific individuals in remote locations through the reproduction of natural dialogue and the accompanying mouth movements of their avatars.

[1327] (Application Example 1)

[1328] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1329] Conventional speech recognition and voice response systems have limitations in achieving natural dialogue with users, and systems combining speech synthesis and lip-syncing have not been widely implemented. Furthermore, in food delivery services, it has been difficult to provide personalized menu suggestions and answer questions through dialogue with users. This invention aims to solve these problems and enable users to use food delivery services more comfortably.

[1330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1331] In this invention, the server includes means for collecting voice samples from a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, virtual assistant means for analyzing the user's questions, generating an appropriate response, and outputting that response using speech synthesis technology, and means for using a generative AI model that generates a response based on the analysis results of the user's voice input. This enables a more natural and intuitive conversation when a user uses a food delivery service.

[1332] A "voice sample" is data that is a recording of a person speaking in a remote location.

[1333] "Speech features" refer to characteristic information such as pronunciation patterns and intonation extracted from collected speech samples.

[1334] A "deep learning model" is a model that uses a multi-layered neural network to learn from large amounts of data and perform predictions and classifications on new data.

[1335] A "virtual assistant" is a software system that generates appropriate responses to user voice input and outputs those responses using speech synthesis technology.

[1336] A "generative AI model" is an artificial intelligence model trained to generate responses based on a user's voice input.

[1337] A "lip-syncing method" is a means of displaying the mouth movements of an avatar in sync with the generated audio.

[1338] A "food delivery service" is a service that delivers food to customers based on their orders.

[1339] This invention is a system that collects voice samples from people in remote locations and uses them to provide users with natural, real-time conversations. This system is particularly effective in food delivery services, where users can use devices such as smartphones or smart glasses to interact with a virtual assistant and obtain information about menus and recommended dishes.

[1340] First, the server collects voice samples and temporarily stores the data. These voice samples are obtained when the user speaks into their smartphone or smart glasses. For example, if the user asks, "What pasta dish do you recommend?", the voice is recorded by the device and sent to the server.

[1341] Next, the server analyzes the collected audio data and converts it into text using speech recognition technology. Simultaneously, it extracts audio features (such as pronunciation patterns and intonation) and stores them in a database. This allows a deep learning model to be trained using the audio features.

[1342] A trained deep learning model generates appropriate responses to voice input from the user. When the user provides voice input again, the server analyzes the speech to understand the intent and uses the generative AI model to generate an appropriate response. This response is then converted into audio data using speech synthesis technology and output to the user.

[1343] Furthermore, lip-sync data synchronized with this generated audio data is created, allowing the virtual assistant avatar to reproduce natural mouth movements. This allows users to visually confirm that the virtual assistant is speaking naturally.

[1344] As a concrete example, when a user asks, "What pasta dish do you recommend?", the system generates a response through the following process: It replies with voice, "My recommendation is carbonara," and an avatar makes the same response with natural mouth movements. This allows the user to experience an immersive conversation.

[1345] The hardware used will be a smartphone, smart glasses, and a server. The software used will be speech recognition technology (e.g., the speech_recognition library), speech synthesis technology (e.g., the gtts library), and a generative AI model (e.g., the GPT-2 model from the transformers library).

[1346] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1347] Step 1:

[1348] When a user speaks to their smartphone or smart glasses, their voice is recorded by the device. The recorded voice data is then sent directly to the server. For example, if a user asks, "What pasta dish do you recommend?", that voice will be collected. The input is the user's voice data, and the output is the voice data sent to the server.

[1349] Step 2:

[1350] The server temporarily stores the received audio data. Then, it converts the audio data into text data using speech recognition technology. Here, the `speech_recognition` library is used to convert speech to text. The input is audio data, and the output is text data.

[1351] Step 3:

[1352] The server further analyzes the text data and extracts speech features (pronunciation patterns, intonation, etc.). A speech feature extraction algorithm is used for the analysis, and the extracted features are stored in a database. The input is text data, and the output is speech feature data.

[1353] Step 4:

[1354] The server accumulates speech feature data and uses it to train a deep learning model. A large amount of speech features are used for training, and a trained generative AI model is constructed. The input is speech feature data, and the output is the trained generative AI model.

[1355] Step 5:

[1356] When the user speaks a question to the device again, the device records the audio and sends it to the server. The server analyzes the received audio, converts it into text data, and performs morphological analysis to understand the user's intent. The input is the newly collected audio data, and the output is the analyzed text data and the results of the intent analysis.

[1357] Step 6:

[1358] The server uses a trained generative AI model to generate appropriate responses to user voice input. The generated responses are converted into speech data using speech synthesis technology (e.g., the GTTS library). The input is the intent analysis result, and the output is the speech response data.

[1359] Step 7:

[1360] The server generates lip-sync data synchronized with the generated audio data, ensuring that the virtual assistant avatar reflects natural mouth movements. The lip-sync data is used to control the avatar's mouth movements in real time. The input is audio data, and the output is lip-sync data.

[1361] Step 8:

[1362] The terminal receives audio data and lip-sync data from the server, plays an audio response to the user, and displays the avatar's lip-sync. The input is audio data and lip-sync data, and the output is the audio response to the user and the avatar's natural mouth movements.

[1363] Through the steps outlined above, users can ask questions about the food delivery service in a natural, conversational format. This allows users to obtain detailed information through dialogue, significantly improving convenience.

[1364] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1365] This invention is a system that reproduces the voice characteristics of a specific person in a remote location, recognizes the user's emotions, and adjusts the response accordingly. The system consists of a smart speaker and a server, and each process is performed as follows.

[1366] 1. Collection and storage of audio samples:

[1367] The user speaks to the smart speaker, which provides a voice sample of a specific person. For example, the user might say "Good morning."

[1368] The device (smart speaker) records the user's voice and sends that voice data to the server.

[1369] 2. Analysis of audio data:

[1370] The server stores the received audio data and converts it into text data using speech recognition technology. Based on the converted text, it extracts audio features such as pronunciation patterns and intonation.

[1371] The extracted speech features are stored in a database and used to train a deep learning model.

[1372] 3. Training deep learning models:

[1373] The server uses the accumulated speech features to train a deep learning model to learn specific speaking styles and intonations.

[1374] The trained model will be used to generate future user responses.

[1375] 4. User emotion recognition:

[1376] When a user speaks a new question or engages in a conversation with the smart speaker, the device records the audio again and sends it to the server.

[1377] The server analyzes the received audio data and uses emotion recognition technology to analyze the user's emotions.

[1378] 5. Analyzing the intent behind questions and conversations:

[1379] The server converts the audio data into text, performs morphological analysis, and understands the user's intent.

[1380] 6. Response generation and adjustment:

[1381] Using a pre-trained deep learning model, it generates appropriate responses to user questions. For example, it can generate a response like, "I had a relaxing day today."

[1382] The response is adjusted based on the user's emotions as assessed by the emotion recognition engine. For example, if the user appears sad, it might generate a response such as, "I had a relaxing day today. You seem a little down, are you okay?"

[1383] The generated text responses are converted into speech data using text-to-speech (TTS) technology.

[1384] 7. Lip sync and avatar display:

[1385] The server generates lip-sync data corresponding to the audio data, naturally expressing the avatar's mouth movements and facial expressions.

[1386] The device receives audio data and lip-sync data, plays the audio, and displays the avatar's facial expressions and mouth movements.

[1387] 8. User experience:

[1388] Users can listen to the responses generated through the smart speaker and see the avatar's natural mouth movements and facial expressions.

[1389] This allows users to enjoy natural conversations with family members or deceased loved ones who live far away, and to receive responses that are sensitive to the user's emotions.

[1390] In this way, the present invention realizes a system that realistically reproduces the voice characteristics of a specific person, understands the user's emotions, and provides a more natural and approachable conversational experience.

[1391] The following describes the processing flow.

[1392] Step 1:

[1393] The user provides a voice sample by speaking to the smart speaker. For example, the user might say "Good morning."

[1394] Step 2:

[1395] The device (smart speaker) records the user's voice and saves it to a temporary buffer. It then sends the recorded voice data to the server.

[1396] Step 3:

[1397] The server receives the audio data sent from the terminal. The received data is saved, and analysis begins.

[1398] Step 4:

[1399] The server converts the audio data into text data using speech recognition technology. Based on the converted text, it extracts audio features (such as pronunciation patterns and intonation).

[1400] Step 5:

[1401] The server stores the extracted speech features in a database. Simultaneously, this data is used to train a deep learning model.

[1402] Step 6:

[1403] The user initiates a new question or conversation with the smart speaker. For example, they might ask, "How was your day?"

[1404] Step 7:

[1405] The device (smart speaker) re-records the user's voice and sends it to the server.

[1406] Step 8:

[1407] The server analyzes the received audio data and converts it back into text data. This text is then subjected to morphological analysis to determine the user's intent.

[1408] Step 9:

[1409] The server uses an emotion recognition engine during the process of analyzing voice data to analyze the user's emotions (such as joy, anger, sadness, etc.).

[1410] Step 10:

[1411] The server evaluates the user's emotions and uses a trained deep learning model to generate an appropriate response. For example, it might generate a response like, "I had a relaxing day."

[1412] Step 11:

[1413] The server adjusts its response based on the emotion recognition results. For example, if the user seems sad, it might change the response to something like, "I had a relaxing day today. You seem a little down, are you okay?"

[1414] Step 12:

[1415] The generated text response is converted into speech data using text-to-speech (TTS) technology. This makes the response sound more natural.

[1416] Step 13:

[1417] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's mouth movements and facial expressions.

[1418] Step 14:

[1419] The server sends audio data and lip-sync data to the terminal.

[1420] Step 15:

[1421] The device (smart speaker) receives audio data and lip-sync data from the server. It plays the audio and displays the avatar's facial expressions and mouth movements.

[1422] Step 16:

[1423] Users listen to the generated responses and observe the avatar's natural mouth movements and facial expressions. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[1424] This specific processing flow allows the present invention to faithfully reproduce the voice characteristics and pronunciation patterns of a particular person and provide a friendly response that corresponds to the user's emotions.

[1425] (Example 2)

[1426] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1427] Providing a natural and engaging conversational experience with a specific person located remotely is challenging. Furthermore, recognizing the user's emotions and generating appropriate responses accordingly is also difficult. Therefore, the challenge lies in accurately reproducing a specific person's voice and speaking style while simultaneously providing responses that resonate with the user's emotions.

[1428] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1429] In this invention, the server includes means for collecting voice samples of a specific person located remotely, means for analyzing the collected voice samples to extract voice features such as pronunciation patterns and intonation, means for training a deep learning model using the voice features, means for recognizing the user's emotions and adjusting the response content, and means for converting voice input from the user into text and analyzing the intent. This makes it possible to realistically reproduce the voice of a specific person and provide a natural conversational experience that responds to the user's emotions.

[1430] A "voice sample" is a digital recording of a user's speech, which imitates the voice of a specific person.

[1431] "Pronunciation patterns" refer to characteristics of speech such as the flow of phonemes, rhythm, stress, and accentuation.

[1432] "Intonation" refers to the changes in pitch and intonation of a voice in spoken language.

[1433] "Speech features" are specific data points extracted from speech samples, and include elements such as pronunciation patterns, intonation, pitch, and length.

[1434] A "deep learning model" is a type of artificial intelligence trained using large amounts of data, and specifically refers to models that use neural networks.

[1435] "Emotion recognition" refers to technology that analyzes and identifies a user's emotional state from their voice or text.

[1436] An "avatar" refers to a representation of a person or character displayed in a virtual space for interaction with a user.

[1437] "Lip sync" refers to a technology that synchronizes the mouth movements of an avatar with the audio.

[1438] "Speech recognition technology" is a technology for converting speech into text, and includes the process of analyzing a speech sample and converting it into text data.

[1439] "Speech synthesis technology" refers to the technology that generates speech based on text data.

[1440] Modes for carrying out the invention

[1441] This invention relates to a system that reproduces the voice characteristics of a specific person located remotely, recognizes the user's emotions, and adjusts its response accordingly. The system consists of a smart speaker and a server and includes the following elements:

[1442] 1. Collection and storage of audio samples

[1443] Voice samples are collected when the user speaks to the smart speaker. For example, the user says "Good morning." The device (smart speaker) records the user's voice with its microphone and temporarily stores this voice data in its internal memory. Next, the device sends this voice data to the server via Wi-Fi.

[1444] 2. Analysis of audio data

[1445] The server stores the received audio data in a dedicated database. Then, it uses speech recognition technology to convert the audio data into text data. Specifically, it uses the Google Speech-to-Text API to convert the audio to text. Based on this text data, it extracts speech features such as pronunciation patterns and intonation, and stores them in the database.

[1446] 3. Training of deep learning models

[1447] A deep learning model is trained using speech features. The server uses a framework such as TensorFlow to learn specific speech patterns and intonations. This training requires a large amount of speech data and corresponding text data. Once training is complete, the model is saved to a specific folder.

[1448] 4. User emotion recognition

[1449] When the user speaks a question or engages in conversation with the smart speaker again, the device records the voice again and sends the audio data to the server. The server analyzes the received audio data and uses emotion recognition technology (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. This analysis result is stored in JSON format.

[1450] 5. Analyzing the intent behind questions and conversations

[1451] The server converts the audio data back into text and performs morphological analysis. For example, MeCab is used for this purpose. Morphological analysis helps understand the user's intent and obtain the information necessary to generate an appropriate response.

[1452] 6. Response generation and adjustment

[1453] Using a pre-trained deep learning model, the server generates appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, considering the user's mood, it might generate additional follow-up responses such as, "You seem a little down, are you okay?" The generated text responses are then converted into speech data using the Google Text-to-Speech API.

[1454] 7. Lip sync and avatar display

[1455] The server generates lip-sync data corresponding to the audio data, creating data to naturally represent the avatar's facial expressions and mouth movements. The terminal receives this audio data and lip-sync data, playing the audio while simultaneously displaying the avatar's facial expressions and mouth movements on the screen.

[1456] 8. User experience

[1457] Users can hear voice responses from smart speakers and see the natural facial expressions and mouth movements of the avatar on the display. This allows them to enjoy a natural conversational experience with family members or deceased loved ones who live far away.

[1458] Specific example

[1459] For example, if a user says "Good morning" to a smart speaker, the smart speaker records the voice and sends it to a server. The server analyzes this voice, extracts features, and stores them in a database. Then, it trains a deep learning model using TensorFlow. If the user asks another question, such as "How was your day?", the server uses emotion recognition technology to analyze the user's emotions. For example, if the user sounds sad, it might generate a response like, "I had a relaxing day. You seem a little down, are you okay?" and convert it into speech using the Google Text-to-Speech API. The device plays this speech, and the display shows the avatar's facial expressions and mouth movements.

[1460] Example of a prompt

[1461] Please generate example responses to the question, "How was your day?" Include additional, emotionally sensitive follow-up if the user seems sad.

[1462] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1463] Step 1: Collect and save audio samples

[1464] The user imitates a specific person's voice and says "Good morning" to the smart speaker. The smart speaker, as the device, records this voice using its microphone. The recorded voice data is temporarily stored in internal memory and then sent to a server via Wi-Fi.

[1465] Input: User's voice "Good morning"

[1466] Output: Recorded audio data (digital format)

[1467] Step 2: Analysis of the voice moon

[1468] The server stores the received audio data and uses speech recognition technology to convert the audio into text data. Using the Google Speech-to-Text API, the audio data is converted into the text "Good morning." Based on this text data, speech features such as pronunciation patterns and intonation are extracted and stored in a new database.

[1469] Input: Recorded audio data

[1470] Output: Text data and speech features

[1471] Step 3: Training the deep learning model

[1472] The server trains a deep learning model using speech features. Libraries such as TensorFlow are used to learn specific speaking styles and intonations. In this training example, a large amount of speech data and corresponding text data are used to enable the model to mimic specific speaking styles and intonations. The trained model is saved to a specific folder.

[1473] Input: Speech features

[1474] Output: Trained deep learning model

[1475] Step 4: User emotion recognition

[1476] The user speaks to the smart speaker again, for example, asking, "How was your day?" The device records this voice and sends it to the server. The server receives the voice data and analyzes the user's emotions using emotion recognition technology (e.g., IBM Watson Tone Analyzer). The analysis results are stored in JSON format.

[1477] Input: User's voice "How was your day?"

[1478] Output: User sentiment analysis results

[1479] Step 5: Analyzing the intent behind questions and conversations

[1480] The server then converts the received audio data back into text data and performs morphological analysis. This is done using a morphological analyzer such as MeCab. By converting the audio into text data such as "How was your day?" and performing morphological analysis, the server understands what the user is asking.

[1481] Input: Audio data

[1482] Output: Analyzed text data and user intent

[1483] Step 6: Generating and adjusting the response

[1484] The server uses a trained deep learning model to generate appropriate responses to user questions. For example, it might generate a text response like, "I had a relaxing day." Furthermore, based on the sentiment recognition results, it might generate a follow-up response such as, "You seem a little down, are you okay?" This generated response is then converted to speech using the Google Text-to-Speech API.

[1485] Input: Analyzed text data and user sentiment analysis results

[1486] Output: Adjusted voice response

[1487] Step 7: Lip sync and avatar display

[1488] The server generates lip-sync data that synchronizes with the generated audio data. This lip-sync data naturally expresses the avatar's mouth movements and facial expressions. The terminal receives the audio data and lip-sync data, plays the audio, and simultaneously displays the avatar on the screen. The avatar's mouth moves in accordance with the lip-sync data, displaying appropriate facial expressions.

[1489] Input: Adjusted voice response

[1490] Output: Avatar facial expressions corresponding to audio and lip-sync data

[1491] Step 8: User Experience

[1492] Users can hear voice responses played from the smart speaker. They can also feel like they are actually having a conversation by observing the natural mouth movements and facial expressions of the avatar displayed on the screen. This system is designed to allow users to enjoy a natural conversational experience with a specific person.

[1493] Input: User's question

[1494] Output: Natural conversation experience and avatar display

[1495] (Application Example 2)

[1496] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1497] A system that can reproduce the voice of a specific person in a remote location and adjust its response by recognizing the user's emotions can provide a natural, emotionally resonant conversational experience in the home and in daily life. However, while such a system could also be applied to work environments such as factories to provide emotional support to employees and improve work efficiency, current systems do not adequately address this need. There is a need for a new system that allows factory workers to relax and boost morale through direct interaction.

[1498] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice samples of a specific person in a remote location, means for analyzing the collected voice samples and extracting voice features, means for training a deep learning model using the voice features, means for generating a response in response to voice input from the user using the trained deep learning model, means for outputting the generated response as voice, means for displaying the mouth movements of an avatar in synchronization with the output voice, and means for recognizing the user's emotions and adjusting the response based on the results. This makes it possible to provide a natural conversational experience that is empathetic to emotions, even in work environments such as factories, enabling mental support for employees and improvement of work efficiency.

[1499] A "voice sample" refers to audio data recorded for use in other systems or algorithms, specifically recording the voice of a particular person.

[1500] "Speech features" refer to characteristic data representations extracted from speech data, such as pronunciation patterns and intonation.

[1501] A "deep learning model" refers to a type of machine learning model that is constructed using multi-layer neural networks to learn complex patterns from data.

[1502] "User voice input" refers to the act of a system user entering information by voice, or the input data itself.

[1503] "Response" refers to the reply or content of the response that the system generates in response to voice input from the user.

[1504] An "avatar" refers to a virtual person or character displayed on a computer screen or by a robot, whose mouth movements and facial expressions are synchronized with the voice.

[1505] "Emotion recognition" refers to a technology or process that analyzes audio data and other information to identify the emotional state of a speaker.

[1506] A "server" refers to a computer system that stores and processes data on a network and communicates with client devices.

[1507] "Means of collection" refers to hardware and software for recording or collecting audio samples via a network.

[1508] "Means of analysis" refers to algorithms and software used to analyze collected data and extract and process speech features.

[1509] "Training methods" refer to the process and techniques of optimizing deep learning models using speech features to improve their skills so that they can accurately generate responses.

[1510] "Means of output" refers to speakers or speech synthesis technology that allow the user to hear the generated response as audio.

[1511] "Means of display" refers to monitors and display technologies used to show users the mouth movements and facial expressions of avatars.

[1512] "Means of adjustment" refers to algorithms and technologies that adjust responses generated based on the results of emotion recognition to create appropriate responses that match the user's emotions.

[1513] This invention provides a system for use in work environments such as factories that reproduces the voice of a specific person, recognizes the emotions of employees, and generates appropriate responses. This can provide emotional support to employees and improve work efficiency.

[1514] The server stores collected audio samples and converts them into text data using speech recognition technology. It also extracts speech features such as pronunciation patterns and intonation, stores them in a database, and uses them to train a deep learning model. The server trains the deep learning model to create a model that learns specific speaking styles and intonations. Then, it uses the trained model to generate appropriate responses to user questions and outputs them as audio. For audio output, text-to-speech (TTS) technology is used to convert the text responses into audio data.

[1515] Furthermore, the server recognizes the user's emotions and adjusts the response generated based on that. For example, if the user appears sad, it generates a response that takes those emotions into consideration. In this way, a natural conversational experience is achieved.

[1516] The terminal (such as a robot or monitor) uses audio and lip-sync data received from the server to play the audio and display the avatar's facial expressions and mouth movements. This allows the user to hear the audio response and see the avatar's natural mouth movements and expressions.

[1517] Specifically, the `speech_recognition` library is used for speech recognition, and the `transformers` library is used for emotion recognition. The `pyttsx3` library is used to convert text to speech.

[1518] The following are specific examples of prompt statements.

[1519] "Please generate an appropriate response based on my emotions. Analyze the following text to recognize my emotions and then provide a friendly voice response accordingly."

[1520] In this way, the present invention provides a natural conversational experience that is attentive to the emotions of employees, even in work environments such as factories, thereby achieving emotional support and improved work efficiency.

[1521] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1522] Step 1:

[1523] The user provides a voice sample. The user speaks to the robot and provides a voice sample of a specific person. For example, the user might say "Good morning." This voice sample is recorded by the device and sent to the server.

[1524] Step 2:

[1525] The server analyzes the audio sample. The server stores the received audio data and converts it into text data using speech recognition technology (e.g., the speech_recognition library). Based on the converted text, it extracts speech features such as pronunciation patterns and intonation. These features are stored in a database.

[1526] Step 3:

[1527] The server trains a deep learning model. The server uses accumulated speech features to train a deep learning model to learn specific speech patterns and intonations. During this process, the transformers library is used to optimize the model using speech features as input. The trained model is then used to generate future user responses.

[1528] Step 4:

[1529] The system recognizes the user's emotions. When the user speaks to the robot again, the terminal records the voice again and sends it to the server. The server analyzes the received voice data and uses emotion recognition technology (e.g., an emotion recognition model from the transformers library) to analyze the user's emotions. The emotion label is output from the voice data as input.

[1530] Step 5:

[1531] The system analyzes the intent behind questions and conversations. The server converts the audio data into text (for example, using the speech_recognition library) and performs morphological analysis to understand the user's intent. This analysis outputs the intent in text format, and a response is determined based on it.

[1532] Step 6:

[1533] The server generates a response and outputs it as speech. A trained deep learning model is used to generate an appropriate response to the user's question. For example, a response such as "I had a relaxing day today" might be generated. The generated text response is converted into speech data using speech synthesis technology (e.g., the pyttsx3 library) and output.

[1534] Step 7:

[1535] The device performs lip-syncing and displays the avatar. The server generates lip-sync data corresponding to the audio data (for example, using an animation library) to create data that naturally expresses the avatar's mouth movements and facial expressions. The device uses this data to play the audio while displaying the avatar's mouth movements and facial expressions.

[1536] Step 8:

[1537] The user experience is enhanced. Users can listen to responses generated through the robot and see the avatar's natural mouth movements and facial expressions. This allows users to enjoy a natural conversational experience with family members or deceased loved ones who are far away, and receive responses that are more emotionally resonant.

[1538] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1539] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1540] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1541] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1542] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1543] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1544] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1545] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1546] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1547] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1548] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1549] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1550] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1551] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1552] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1553] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1554] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1555] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1556] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1557] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1558] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1559] The following is further disclosed regarding the embodiments described above.

[1560] (Claim 1)

[1561] A means of collecting voice samples from a specific person in a remote location,

[1562] A means for analyzing collected audio samples and extracting audio features,

[1563] A method for training a deep learning model using speech features,

[1564] A means for generating a response in response to voice input from a user using a trained deep learning model,

[1565] A means for outputting the generated response as audio,

[1566] A means of displaying the avatar's mouth movements in sync with the outputted audio,

[1567] A system that includes this.

[1568] (Claim 2)

[1569] The system according to claim 1, which extracts the pronunciation patterns and intonation of a specific person located in a remote location as features.

[1570] (Claim 3)

[1571] The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent.

[1572] "Example 1"

[1573] (Claim 1)

[1574] Means of acquisition,

[1575] A means of saving the acquired audio data,

[1576] A means of converting stored audio data into text using speech recognition technology,

[1577] A means of extracting and storing speech features,

[1578] A method for training a deep learning model using accumulated speech features,

[1579] A means for generating a response in response to voice input from a user using a trained deep learning model,

[1580] A means for outputting the generated response as speech using speech synthesis technology,

[1581] A means of displaying the avatar's mouth movements in sync with the outputted audio,

[1582] A system that includes this.

[1583] (Claim 2)

[1584] The system according to claim 1, which extracts the pronunciation patterns and intonation of a specific person located in a remote location as features.

[1585] (Claim 3)

[1586] The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent.

[1587] "Application Example 1"

[1588] (Claim 1)

[1589] A means of collecting voice samples from a specific person in a remote location,

[1590] A means for analyzing collected audio samples and extracting audio features,

[1591] A method for training a deep learning model using speech features,

[1592] A means for generating a response in response to voice input from a user using a trained deep learning model,

[1593] A means for outputting the generated response as audio,

[1594] A means of displaying the avatar's mouth movements in sync with the outputted audio,

[1595] A virtual assistant means that analyzes user questions, generates appropriate responses, and outputs those responses using speech synthesis technology,

[1596] A means of using a generative AI model that generates a response based on the analysis results of the user's voice input,

[1597] A system that includes this.

[1598] (Claim 2)

[1599] The system according to claim 1, which extracts the pronunciation patterns and intonation of a specific person located in a remote location as features.

[1600] (Claim 3)

[1601] The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent.

[1602] "Example 2 of combining an emotion engine"

[1603] (Claim 1)

[1604] A means of collecting voice samples from a specific person in a remote location,

[1605] A method for analyzing collected audio samples to extract speech features such as pronunciation patterns and intonation,

[1606] A method for training a deep learning model using speech features,

[1607] A means for generating a response in response to voice input from a user using a trained deep learning model,

[1608] A means of recognizing the user's emotions and adjusting the response accordingly,

[1609] A means for outputting the generated response as audio,

[1610] A means of displaying the avatar's mouth movements in sync with the outputted audio,

[1611] A system that includes this.

[1612] (Claim 2)

[1613] The system according to claim 1, which extracts pronunciation patterns and intonation as speech features.

[1614] (Claim 3)

[1615] The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent.

[1616] "Application example 2 when combining with an emotional engine"

[1617] (Claim 1)

[1618] A means of collecting voice samples from a specific person in a remote location,

[1619] A means for analyzing collected audio samples and extracting audio features,

[1620] A method for training a deep learning model using speech features,

[1621] A means for generating a response in response to voice input from a user using a trained deep learning model,

[1622] A means for outputting the generated response as audio,

[1623] A means of displaying the avatar's mouth movements in sync with the outputted audio,

[1624] A means of recognizing the user's emotions and adjusting the response based on the results,

[1625] A system that includes this.

[1626] (Claim 2)

[1627] The system according to claim 1, which extracts the pronunciation patterns and intonation of a specific person located in a remote location as features.

[1628] (Claim 3)

[1629] The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent. [Explanation of Symbols]

[1630] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting voice samples from a specific person in a remote location, A means for analyzing collected audio samples and extracting audio features, A method for training a deep learning model using speech features, A means for generating a response in response to voice input from a user using a trained deep learning model, A means for outputting the generated response as audio, A means of displaying the avatar's mouth movements in sync with the outputted audio, A system that includes this.

2. The system according to claim 1, which extracts the pronunciation patterns and intonation of a specific person located in a remote location as features.

3. The system according to claim 1, comprising means for converting voice input from a user into text and analyzing the intent.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A